Performance Benchmarking with Criterion
Overview
This document provides a comprehensive guide to performance benchmarking in readstat-rs using Criterion.rs.
Quick Start
# Run all benchmarks (from the repository root)
cargo bench -p readstat
# View HTML reports (Criterion writes to the workspace-root target/)
open target/criterion/report/index.html
What Gets Benchmarked
1. Reading Performance
- Metadata Reading (
~300-950 µs) - File header parsing - Single Chunk Reading - Full dataset read performance
- Chunked Reading - Streaming with different chunk sizes (1K, 5K, 10K rows)
2. Data Conversion
- Arrow Conversion - SAS types → Arrow RecordBatch overhead
3. Writing Performance
- CSV Writing - Text format output
- Parquet Compression - Uncompressed, Snappy, Zstd comparison
- Format Comparison - CSV vs Parquet vs Feather vs NDJSON
4. Parallel Write Optimization
- Row-Group Sizes - Native parallel Parquet column encoding with different rows per group
5. End-to-End Pipeline
- Complete Conversion - Read + Write combined (most important)
Sample Results
From initial benchmark run (example output):
metadata_reading/all_types.sas7bdat
time: [299.41 µs 301.84 µs 304.29 µs]
metadata_reading/cars.sas7bdat
time: [935.21 µs 943.52 µs 952.41 µs]
read_single_chunk/cars.sas7bdat
time: [~2-3 ms]
thrpt: [~150-200K rows/sec]
write_parquet_compression/snappy
time: [~4-6 ms]
thrpt: [~70-100K rows/sec]
end_to_end_conversion/parquet
time: [~6-9 ms]
thrpt: [~50-70K rows/sec]
Interpreting Results
Understanding the Output
Time Measurement:
time: [299.41 µs 301.84 µs 304.29 µs]
^ ^ ^
| | +-- Upper bound (95% confidence)
| +------------ Median
+---------------------- Lower bound (95% confidence)
Throughput:
thrpt: [150K elem/s 175K elem/s 200K elem/s]
^ ^ ^
| | +-- Upper bound
| +-------------- Median
+-------------------------- Lower bound
Change Detection:
change: [-2.3456% -1.2345% +0.1234%] (p = 0.12 > 0.05)
^ ^ ^ ^
| | | +-- Statistical significance
| | +----------- Upper bound of change
| +--------------------- Median change
+------------------------------- Lower bound of change
What to Look For
🔴 Red Flags (Investigate)
- High variance (>10%) - Results unreliable
- Significant regression (>5% slower, p < 0.05)
- Outliers (>5% of samples)
🟡 Opportunities
- Chunked reading - Test if different chunk size improves throughput
- Buffer sizes - If small buffer performs as well as large, save memory
- Compression - If uncompressed only slightly faster, use compression
🟢 Validation
- Low variance (<5%) - Reliable results
- Improvements (>10% faster, p < 0.05)
- Expected patterns (e.g., compression should be slower but smaller)
Performance Optimization Workflow
Step 1: Establish Baseline
# Save current performance as baseline
cargo bench --save-baseline main
# Results saved to target/criterion/{benchmark}/main/
Step 2: Make Changes
Edit code with optimization hypothesis:
- Increase buffer size
- Change algorithm
- Add caching
- Parallel processing
Step 3: Measure Impact
# Compare against baseline
cargo bench --baseline main
# Look for "change: [X% Y% Z%]" in output
Step 4: Analyze & Iterate
If improved (>10%, p < 0.05):
✅ Keep the change
✅ Update baseline: cargo bench --save-baseline main
If no change (<5%): ⚠️ Optimization didn’t help - profile to find real bottleneck
If regressed (slower): ❌ Revert change ❌ Investigate why performance decreased
Common Optimization Scenarios
Scenario 1: Slow Reading
Symptoms: read_single_chunk time is high
Investigate:
- ReadStat C library overhead (FFI calls)
- Memory allocation patterns
- Callback overhead
Try:
- Larger buffers in C library
- Memory-mapped files (see evaluation doc)
- Pre-allocate column vectors
Scenario 2: Slow Writing
Symptoms: write_formats time is high
Investigate:
- BufWriter buffer size
- Format-specific overhead
- Compression CPU usage
Try:
- Increase BufWriter capacity (currently 8KB)
- Use faster compression (Snappy vs Zstd)
- Parallel writing (already implemented)
Scenario 3: Memory Issues
Symptoms: System swapping, OOM errors
Investigate:
- Chunk size too large
- Too many parallel streams
- Memory leaks
Try:
- Reduce
stream_rows(default 10,000) - Reduce parallel write buffer (default 100MB)
- Use bounded channels (already implemented)
Scenario 4: High Variance
Symptoms: Large confidence intervals, many outliers
Investigate:
- System background activity
- CPU frequency scaling
- Thermal throttling
Try:
- Close background apps
- Disable frequency scaling
- Run on consistent power mode
Advanced Profiling
CPU Profiling with Flamegraphs
# Install flamegraph
cargo install flamegraph
# Profile a specific benchmark
cargo flamegraph --bench readstat_benchmarks -- --bench read_single_chunk
# Open flamegraph.svg to see hotspots
What to look for:
- Wide bars = lots of time spent
- Deep stacks = call overhead
- Unexpected functions = bugs/inefficiency
Memory Profiling
# Using valgrind (Linux)
valgrind --tool=massif \
cargo bench read_single_chunk --no-run
ms_print massif.out.* > memory_profile.txt
# Using heaptrack (Linux)
heaptrack cargo bench read_single_chunk
heaptrack_gui heaptrack.*.gz
System Call Tracing
# Linux: strace
strace -c cargo bench read_single_chunk 2>&1 | tail -20
# macOS: dtruss
sudo dtruss -c cargo bench read_single_chunk
Comparing Implementations
Before/After Memory-Mapped Files
# Baseline without mmap
git checkout main
cargo bench --save-baseline without-mmap
# With mmap implementation
git checkout feature/mmap
cargo bench --baseline without-mmap
# Look for improvements in read_single_chunk
Parallel vs Sequential
# Test with different parallelism settings
cargo bench end_to_end -- --parallel
cargo bench end_to_end -- --sequential
CI/CD Integration
Performance Regression Detection
Add to .github/workflows/benchmarks.yml:
name: Performance Benchmarks
on:
pull_request:
branches: [main]
jobs:
benchmark:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Rust
uses: dtolnay/rust-toolchain@stable
- name: Run benchmarks
run: |
cd crates/readstat
cargo bench --no-run # Just compile for CI
- name: Compare with baseline (on main branch)
if: github.event_name == 'pull_request'
run: |
git fetch origin main:main
git checkout main
cargo bench --save-baseline main
git checkout -
cargo bench --baseline main
Best Practices
Do’s ✅
- Run benchmarks on consistent hardware
- Close background applications
- Use
--save-baselinefor comparisons - Profile after benchmarking to find bottlenecks
- Document performance changes in PRs
- Test on representative data sizes
Don’ts ❌
- Don’t benchmark on laptop (throttling)
- Don’t optimize without profiling first
- Don’t trust results with high variance
- Don’t compare across different systems
- Don’t commit benchmark artifacts
- Don’t skip statistical significance checks
Performance Goals
Current Performance (Baseline)
- Metadata reading: ~300-950 µs
- Read throughput: ~150-200K rows/sec
- Write throughput: ~70-100K rows/sec
- End-to-end: ~50-70K rows/sec
Target Performance (Goals)
- Metadata reading: <500 µs (↓30%)
- Read throughput: >250K rows/sec (↑25%)
- Write throughput: >100K rows/sec (↑30%)
- End-to-end: >100K rows/sec (↑40%)
Stretch Goals
- Memory-mapped reads: 2x faster for large files
- Parallel writes: 3-4x speedup with 4+ cores
- Compression: <10% overhead for Snappy
Data Files for Benchmarking
Current Test Data
- all_types.sas7bdat - 3 rows, 10 vars (tiny)
- cars.sas7bdat - 1081 rows, 13 vars (small)
Recommended Additional Data
For comprehensive benchmarking, consider adding:
Small (good for quick iteration):
- < 1 MB file size
- < 1,000 rows
- 5-10 variables
Medium (typical use case):
- 10-100 MB file size
- 10,000-100,000 rows
- 10-50 variables
Large (stress test):
-
1 GB file size
-
1,000,000 rows
- 50+ variables
Resources
Documentation
Tools
- cargo-flamegraph
- cargo-benchcmp
- hyperfine - CLI benchmarking (see below)
Blog Posts
Next Steps
- Run full benchmark suite:
cargo bench - Review HTML reports: Open
target/criterion/report/index.html - Identify bottlenecks: Look for slowest operations
- Profile with flamegraph: Focus on hotspots
- Implement optimizations: Test one at a time
- Validate improvements: Compare against baseline
- Document findings: Update this file with results
Questions?
- See detailed README:
crates/readstat/benches/README.md - Check Criterion docs: https://bheisler.github.io/criterion.rs/book/
- Review performance evaluation: Memory-mapped files analysis (separate doc)
Benchmarking with hyperfine
Benchmarking performed with hyperfine.
For the canonical 4-million-row conversion benchmark, use the repository script:
./scripts/benchmark-conversion.sh --quick # smoke run
./scripts/benchmark-conversion.sh # full writer, CPU, and batch-size sweeps
It downloads and verifies the immutable benchmark corpus, records the machine
and toolchain configuration, builds the release CLI, compares serial and native
parallel Parquet encoding, measures peak memory, sweeps Rayon worker counts and
input batch sizes, and writes JSON plus Markdown reports under
target/benchmark-results/<timestamp>/.
2026-07-27 conversion baseline
Commit 487547285c61cf5aa412e915796fdaa0875e6c40 was measured on an 18-core
Apple Silicon Mac with 64 GiB of memory. On the canonical 4-million-row corpus,
native parallel Parquet writing was 1.30x faster than serial writing (1.511 s
versus 1.966 s) and used 201 MB rather than 553 MB peak RSS. Four Rayon workers
were nominally fastest, but 4 through 18 workers were within approximately 2%.
Input batches from 10,000 through 100,000 rows were also effectively tied; 5,000
rows was 14% slower. The one-pass reader was 2.36x faster than the removed
partitioned reader.
The real Census AHS 2021 household.sas7bdat workload (64,141 rows, 1,078
columns) reversed the memory result: serial writing used 757 MB peak RSS,
parallel 100,000-row groups used 1.38 GB, and parallel 25,000-row groups used
1.15 GB. Smaller groups also made every tested Parquet output 6-7% larger, so
the 100,000-row parallel default remains unchanged. Prefer --serial-write and a
smaller --stream-rows value for unusually wide data.
Format baselines on the canonical corpus were 1.446 s for source-only parsing, 1.431 s for Feather, 3.245 s for NDJSON, and 3.416 s for CSV. Feather remains serial because its writer is already hidden behind parsing and Arrow IPC has no public zero-reencode file assembly API comparable to Parquet column-chunk append. Bounded four-batch parallel text encoding reduced CSV from 3.393 s to 1.551 s (2.19x) and NDJSON from 3.247 s to 1.564 s (2.08x). Complete serial and parallel outputs were byte-identical. Peak RSS increased from 61 MB to 90 MB for CSV and from 62 MB to 99 MB for NDJSON. Four workers saturated both formats; eight provided no further improvement.
This example compares the performance of the Rust binary with the performance of the C binary built from the ReadStat repository. In general, hope that performance is fairly close to that of the C binary.
To run, execute the following from within the readstat directory.
# Windows
hyperfine --warmup 5 "ReadStat_App.exe -f crates\readstat-tests\tests\data\cars.sas7bdat tests\data\cars_c.csv" ".\target\release\readstat.exe data crates\readstat-tests\tests\data\cars.sas7bdat --output crates\readstat-tests\tests\data\cars_rust.csv"
📝 First experiments on Windows are challenging to interpret due to file caching. Need further research into utilizing the --prepare option provided by hyperfine on Windows.
# Linux and macOS
hyperfine --prepare "sync; echo 3 | sudo tee /proc/sys/vm/drop_caches" "readstat -f crates/readstat-tests/tests/data/cars.sas7bdat crates/readstat-tests/tests/data/cars_c.csv" "./target/release/readstat convert crates/readstat-tests/tests/data/cars.sas7bdat --output crates/readstat-tests/tests/data/cars_rust.csv"
Other, future, benchmarking may be performed now that channels and threads have been developed.
Profiling with Flamegraphs
Profiling performed with cargo flamegraph.
The readstat binary lives in the readstat-cli crate, so target it with -p readstat-cli. Run the following from the repository root.
cargo flamegraph -p readstat-cli --bin readstat -- data crates/readstat-tests/tests/data/_ahs2019n.sas7bdat --output crates/readstat-tests/tests/data/_ahs2019n.csv
Flamegraph is written to flamegraph.svg in the directory you run the command from (the repository root).
📝 Have yet to utilize flamegraphs in order to improve performance.
Large external SAS7BDAT baseline (Stage 0)
The large_sas_benchmark example is an argument-driven baseline for the
current high-level ReadStatReader path. It streams batches through visit,
counts their rows, and immediately drops them without writing output. Its timed
region includes the reader’s metadata parse and all data parses, but excludes
argument parsing and the filesystem metadata lookup.
Census AHS 2021 corpus
From the repository root, download and extract the public U.S. Census 2021
American Housing Survey National PUF SAS archive into the ignored target/
tree (never commit this corpus):
mkdir -p target/benchmark-data/ahs-2021 && \
curl --fail --location --output target/benchmark-data/ahs-2021/ahs-2021-sas.zip \
'https://www2.census.gov/programs-surveys/ahs/2021/AHS%202021%20National%20PUF%20v1.0%20SAS.zip' && \
unzip -o target/benchmark-data/ahs-2021/ahs-2021-sas.zip \
-d target/benchmark-data/ahs-2021
The ZIP is approximately 160 MB and contains household.sas7bdat
(approximately 311,689,216 bytes), mortgage.sas7bdat, person.sas7bdat, and
project.sas7bdat. Confirm actual sizes with the harness rather than treating
the approximate published sizes as checksums. If extraction creates a nested
directory, find the input with:
find target/benchmark-data/ahs-2021 -name household.sas7bdat -print
Running and comparing chunk sizes
cargo run --release -p readstat --example large_sas_benchmark -- \
target/benchmark-data/ahs-2021/household.sas7bdat --chunk-rows 10000
for rows in 1000 10000 100000; do
cargo run --quiet --release -p readstat --example large_sas_benchmark -- \
target/benchmark-data/ahs-2021/household.sas7bdat --chunk-rows "$rows"
done
Each run reports source bytes, exact emitted rows and batches, elapsed wall
time, rows/s, source MiB/s, and expected parser invocations. The default
one-pass mode uses one metadata parse plus one data parse regardless of batch
count. legacy-chunked preserves the former parser-per-batch behavior as a
benchmark baseline; its expected parser count is 1 + batches. Both counts are
derived rather than instrumented.
Compare the Stage 1 one-pass reader to the former implementation with identical batch sizing:
for mode in one-pass legacy-chunked; do
cargo run --quiet --release -p readstat --example large_sas_benchmark -- \
target/benchmark-data/ahs-2021/household.sas7bdat \
--chunk-rows 10000 --mode "$mode"
done
On Linux with /proc mounted, current RSS is /proc/self/status VmRSS and
process peak RSS is VmHWM, both in KiB. Other platforms and Linux containers
without /proc report RSS as unavailable. The process high-water mark includes
runtime allocations and is not solely Arrow batch memory.
For reproducible comparisons:
- Record the commit, release profile, CPU, OS, storage type, and exact command.
- Repeat each size and rotate their order. The harness measures one pass rather than providing a statistical framework.
- Label cache state. The first read may be cold-cache and storage-bound; later reads are normally warm-cache due to the OS page cache. Dropping Linux caches requires privileges and affects the whole host, so do not do it on shared systems.
- Avoid concurrent heavy I/O and CPU work; frequency scaling, thermal throttling, and network filesystems can materially affect results.
- MiB/s divides source file bytes by elapsed time. It is workload throughput,
not measured physical I/O:
legacy-chunkedrereads prefixes, while OS caching can serve either mode without issuing storage reads for every source byte.
Synthetic SAS benchmark corpus
The public Census file is a useful real workload, but its household.sas7bdat
member is unusually wide and has only 64,141 rows. The fixed-seed
create_rand_ds.sas program
defines a complementary canonical profile:
- Dataset:
readstat_benchmark_v1.sas7bdat - Seed:
20260727 - Rows: 4,000,000
- Numeric columns: 12 (SAS numerics occupy 8 bytes each)
- Character columns: 8, each 32 bytes wide
- Compression: none
- Expected raw row payload: 352 bytes, excluding SAS page overhead
- Expected raw payload total: 1,408,000,000 bytes (approximately 1.31 GiB)
This shape is deliberately tall enough to reveal repeated-prefix parsing and
reader partitioning costs. Its high-entropy printable strings also exercise
string conversion and avoid making compression ratio the dominant variable.
The SAS version, host platform, session encoding, and page settings can affect
the random-number implementation or binary representation even with a fixed
seed. The generator writes these details, its parameters, and the complete
PROC CONTENTS listing to $HOME/readstat_benchmark_v1_manifest.txt in the
same run. On Linux, a SAS session with the XCMD option also appends ls -lh
and sha256sum output through FILENAME PIPE. A restricted NOXCMD session
records that limitation; download the dataset and manifest together, then let
the publication script calculate and verify the size and SHA-256 locally.
The manifest normally records the output size and digest automatically. With
NOXCMD, the same information can be calculated after downloading. On macOS,
du reports allocated disk usage while stat reports the exact logical size
used for GitHub’s asset limit:
ls -lh readstat_benchmark_v1.sas7bdat # human-readable logical size
du -h readstat_benchmark_v1.sas7bdat # allocated disk usage
stat -f '%z bytes' readstat_benchmark_v1.sas7bdat
shasum -a 256 readstat_benchmark_v1.sas7bdat
Then validate the file with both benchmark modes and several batch sizes before publishing it. Exact row counts must agree in every run.
Do not commit the generated file or add it to Git LFS. Git LFS charges the
repository owner for stored versions and download bandwidth, making a frequently
downloaded benchmark corpus an unnecessary repository cost. Prefer a dedicated
GitHub release such as benchmark-data-v1; GitHub release assets do not consume
Git history and have no aggregate size or bandwidth limit. Each individual
release asset must remain under 2 GiB. If the generated SAS file exceeds that
limit, reduce the canonical row count rather than splitting the file: a split
archive is awkward for automated benchmark setup and obscures the actual source
size.
Publish these alongside the SAS file:
- A SHA-256 checksum file.
- The exact generator program or its repository commit.
- The generated
readstat_benchmark_v1_manifest.txtfile.
The publication script validates all three, confirms the dataset has exactly
4,000,000 readable rows, checks the 2 GiB asset limit and current origin/main,
and refuses to replace an existing benchmark release. Run it without arguments
for a non-destructive preview, then opt in to publication explicitly. First
download the SAS file and manifest into the repository’s ignored local data
directory:
benchmark-data/readstat_benchmark_v1.sas7bdat
benchmark-data/readstat_benchmark_v1_manifest.txt
Then run:
./scripts/publish-benchmark.sh
./scripts/publish-benchmark.sh --publish
It creates the immutable benchmark-data-v1 tag and release, uploads the SAS
file, manifest, and generated checksum sidecar, and marks the release as not
“Latest” so it does not displace the current software release. Override the
repository-local paths only when necessary with BENCHMARK_DATASET and
BENCHMARK_MANIFEST. The script uses sha256sum on Linux or shasum -a 256
on macOS. The generated files are ignored by Git and must never be force-added
or placed in Git LFS.
The Census and synthetic datasets answer different questions and should both be retained in benchmark reports; the synthetic corpus must not replace validation against real files.