Skip to content

test(benchmark): codec vs per-read cost decomposition (stacked on #9) - #10

Draft
dougnukem wants to merge 4 commits into
benchmark/full-size-profilingfrom
bench/codec-io-decomposition
Draft

dougnukem wants to merge 4 commits into
benchmark/full-size-profilingfrom
bench/codec-io-decomposition

Conversation

@dougnukem

@dougnukem dougnukem commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #9. Splits fastp's processing time by kind of work: gzip decompression, parsing and per-read QC, gzip compression, and writing. #9 splits it by stage in time. This split decides which optimisations can pay off.

What it adds (benchmark/codec-io.md has the method, how to read it, and the results):

  • scripts/io_variants.sh: the same binary on the same reads with codec stages switched off. Inputs are gz, plain, or BGZF; outputs are gz, plain, or none. plain_none is the per-read compute floor. The one-off cost of making the plain and BGZF inputs goes to prep.tsv.
  • scripts/codec_baselines.sh: standalone codec MB/s per thread count (ISA-L, libdeflate levels 1/4/6, pigz, rapidgzip, bgzip).
  • scripts/threadcpu.py: per-thread CPU sampler, opt-in from bench_full.py via FASTP_BENCH_THREADS_JSON.
  • scripts/decompose.py: per-stage wall and CPU differences, the single-thread reader ceiling, busiest threads, and a check that output digests don't differ across I/O modes.
  • bench_full.py: digests plain outputs and allows runs with no -o. TSV columns unchanged.

Results (n2d-highmem-48, 4 complete public datasets, -w 16/48, 9 variants × 2 reps; data in benchmark/results/codec-io-2026-09/):

  • Per-read work is ~75% of CPU on every dataset. Output compression is 17–25%, input decompression 1–5%. Removing both codecs saves only 13–26% of wall.
  • -w 48 is slower than -w 16 on 3 of 4 datasets (atac 200 → 270 s) and costs 25–45% more CPU. The box is 24 cores × 2 SMT threads.
  • BGZF input doesn't help. The parse stays on the reader, and converting to BGZF costs ~5.7 µs of CPU per read.
  • Output digests are identical across all 72 variant/thread cells.
  • Uncapped parallel inflate: rapidgzip 3.4 GB/s at 16 threads (~3.7× ISA-L's CPU per byte); bgzip on BGZF 9.0 GB/s.

Source facts (in codec-io.md):

  • Ordinary .gz input is inflated and parsed by one thread per mate.
  • Output compression is already parallel.
  • The BGZF pool is max(1, (nproc − (-w+4)) / gz_inputs), which is 1 thread at -w ≥ 44 on 48 vCPUs.

Self-review fixes already in:

  • threadcpu role grouping (fp-write/fp-bgzf were merged into one role).
  • Decompression baselines write to /dev/null; piping through wc capped them at ~2.5 GB/s.

🤖 Generated with Claude Code

Splits fastp's processing time by kind of work (gzip decompression, parsing and
per-read QC, gzip compression, writing) by running the same binary on the same
reads with codec stages switched off: gz/plain/BGZF input x gz/plain/no output.
Adds standalone codec baselines (ISA-L, libdeflate, pigz, rapidgzip, bgzip) as
ceilings, a per-thread CPU sampler (threadcpu.py, opt-in from bench_full.py) to
show a pinned reader thread, and decompose.py to attribute wall and CPU per stage.

bench_full.py: digest plain (non-.gz) outputs with cat, and allow runs with no -o.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rsplit('-') merged fp-write and fp-bgzf into 'fp' and fp-read-L/R into one role.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Piping decompressed output through wc -c capped parallel decompressors (bgzip,
rapidgzip) at ~2.5 GB/s on n2d-highmem-48; the uncompressed size is already known.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Four complete public datasets at -w 16/48, 9 I/O variants, 2 reps: per-read
work is ~75% of CPU, codecs 20-30%; -w 48 is slower than -w 16 on 3 of 4
datasets; BGZF input does not help. Output digests identical across all 72
variant/thread cells.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants