Skip to content

test(benchmark): transform costs + GPU benchmarks for a reformulation cost model (stacked on #11) - #12

Draft
dougnukem wants to merge 9 commits into
bench/pack-tracefrom
bench/reformulation-costs
Draft

dougnukem wants to merge 9 commits into
bench/pack-tracefrom
bench/reformulation-costs

Conversation

@dougnukem

@dougnukem dougnukem commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #11. Asks whether reformulating fastp for a GPU (or other batch accelerator) pays off, counting the cost of transforming data into and out of the accelerator's format. Write-up with results: benchmark/reformulations.md. Data: benchmark/results/reformulation-2026-09/.

Tools

  • tools/soa_pack.cpp: CPU cost of index, pack (SoA / 2-bit) and emit, plus a representative QC kernel, on real reads. --dump writes a batch for the GPU check.
  • tools/gpu_bench.py: PCIe, the same kernel on the GPU (checked read for read against the CPU), nvCOMP batched BGZF inflate, single-stream gzip (nvlzcat lookahead, nvCOMP ≥ 5.2), and GPU gzip compression.
  • scripts/reformulation_model.py: projected CPU, GPU and wall time per reformulation from measured inputs. CPU: D1 BGZF, D2 parallel gzip reader, D3 -z 1, D4 uncompressed streaming, D5 flat-batch parser. GPU: A offload QC, B GPU-resident, C fused into a GPU aligner. QC-speedup sensitivity rows included.
  • scripts/run_reformulation_suite.sh: every CPU-side measurement in one command.

Measured (n2d-highmem-48 + one L4, nvCOMP 5.3):

  • Transform costs are small: index + pack + emit is 80–120 ns/read, against fastp's 670–1100 ns parse. 2-bit packing isn't worth it (4–5× the cost, ~30% fewer bytes).
  • L4 QC kernel: 5.6 ns/read with 0 mismatches vs CPU over 1M reads, but moving the batch over PCIe costs 24.5 ns/read, 4× the kernel.
  • L4 single-stream gzip inflate (nvlzcat): 1.26 GB/s. That's below 16 CPU threads of rapidgzip (3.4 GB/s). The nvCOMP Python default is 31 MB/s.
  • L4 BGZF inflate 7.3 GB/s vs bgzip at 16 CPU threads 9.0 GB/s.
  • L4 gzip at ratio 4.61 runs at 2.1 GB/s. At fastp's -z 4 ratio (4.88) it runs at 0.12 GB/s, about one CPU core.

Conclusion (details in reformulations.md):

  • Largest wins are CPU changes: fix and speed up trimBySequence's gap search (33% of worker CPU, plus a correctness bug), a cheaper parser (1.2–2.1× on reader-bound data), right-sized -w, and uncompressed streaming to the aligner.
  • A pays off for CPU-hours (−70–75% CPU) only if fastp's real per-read work runs ≥ ~10–15× a CPU thread on the GPU. That's unmeasured: the analogue kernel covers ~1/8 of the work.
  • B is limited by single-stream inflate and GPU gzip's ratio.
  • C is the one real GPU opportunity, but it needs a pre-alignment hook that Parabricks doesn't have.

Self-review fixes already in:

  • nvCOMP Python API usage: enums, device sizes, freeing GPU memory, and the non-lookahead gzip path measured on a 100 MB member.
  • soa_pack truncated-record over-read.
  • B uses only the single-stream rate for ordinary gzip.
  • SE BGZF pool formula.
  • Python-API encode rows excluded from the model.

Remaining

🤖 Generated with Claude Code

…n cost model

soa_pack.cpp measures on real reads what a batch-accelerator path adds on the
CPU: record indexing, packing into a fixed-stride SoA batch (and 2-bit),
serialising kept reads back to FASTQ from trim decisions, and a representative
per-read QC kernel (N/low-quality filters, polyG, adapter search with
mismatches, per-position stats). It can dump a batch plus the CPU kernel's
decisions.

gpu_bench.py measures the GPU side: PCIe bandwidth, the same kernel on the
GPU checked read-for-read against the CPU decisions, nvCOMP batched BGZF
inflate, single-stream gzip inflate (nvlzcat lookahead, nvCOMP >= 5.2), and
GPU Deflate/gzip compression.

reformulation_model.py combines trace per-read costs, transform costs and GPU
rates into projected CPU/GPU/wall time (and cost, given prices) for: BGZF
input, parallel gzip reader, -z 1, offloading per-read QC, a GPU-resident
pipeline, and fusing QC into a GPU aligner, with QC-speedup sensitivity rows.
run_reformulation_suite.sh runs all CPU-side measurements in one go.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… to model

BitstreamKind enum instead of a string, uncomp_chunk_size, RAW bitstream for
Gzip/Deflate, algorithm_type sweep, median of 3 nvlzcat runs. Model gains D4:
stream uncompressed output into the next tool.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- gpu_bench.py: sizes of device arrays via cupy (implicit host conversion is
  refused, which made every nvCOMP section record an error); free the memory
  pool between sections and before nvlzcat subprocesses (L4 OOM).
- soa_pack.cpp: skip records whose quality line length differs from the
  sequence (a truncated last record read past the buffer).
- reformulation_model.py: B uses the single-stream gzip rate unless
  --bgzf-input; picks the fastest gzip-compatible GPU compressor within 90% of
  libdeflate's ratio; SE BGZF pool formula (nproc - w - 3); warn without
  --dataset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Python API has no choice of gzip decompression algorithm, and its default
decodes one single-member stream as one chunk: on the whole 2 GB file it ran
for over 15 minutes. nvlzcat (lookahead algorithm) covers the full file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
They run ~1000x slower than nvlzcat on the same data (API path, not GPU
capability), so they must not be picked as the GPU compressor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Option A (offload QC to a GPU) also replaces fastp's Read-object parse with an
index+pack parse. D5 is that parser change alone, so A's GPU contribution can
be read as A vs D5 rather than A vs base.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Survey of existing GPU/CPU work (checked against primary sources), the ways to
reformulate fastp (CPU D1-D7, GPU A/B/C) with their data-transform costs, and
measured results: per-read costs from the traces, transform costs from
soa_pack on 4 datasets, and L4 numbers for PCIe, the QC kernel (0 mismatches
vs CPU over 1M reads), single-stream and BGZF inflate, and GPU gzip ratio vs
throughput. Model projections per dataset at -w 16/48 with QC-speedup
sensitivity. Conclusion: the largest wins are CPU changes (fix and speed up
trimBySequence's gap search, cheaper parsing, right-sized -w, uncompressed
streaming); a GPU pays off only fused into a GPU pipeline, or for CPU-hours if
fastp's real per-read work reaches ~10-15x per thread on the GPU.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants