test(benchmark): transform costs + GPU benchmarks for a reformulation cost model (stacked on #11) - #12
Draft
dougnukem wants to merge 9 commits into
Draft
test(benchmark): transform costs + GPU benchmarks for a reformulation cost model (stacked on #11)#12dougnukem wants to merge 9 commits into
dougnukem wants to merge 9 commits into
Conversation
…n cost model soa_pack.cpp measures on real reads what a batch-accelerator path adds on the CPU: record indexing, packing into a fixed-stride SoA batch (and 2-bit), serialising kept reads back to FASTQ from trim decisions, and a representative per-read QC kernel (N/low-quality filters, polyG, adapter search with mismatches, per-position stats). It can dump a batch plus the CPU kernel's decisions. gpu_bench.py measures the GPU side: PCIe bandwidth, the same kernel on the GPU checked read-for-read against the CPU decisions, nvCOMP batched BGZF inflate, single-stream gzip inflate (nvlzcat lookahead, nvCOMP >= 5.2), and GPU Deflate/gzip compression. reformulation_model.py combines trace per-read costs, transform costs and GPU rates into projected CPU/GPU/wall time (and cost, given prices) for: BGZF input, parallel gzip reader, -z 1, offloading per-read QC, a GPU-resident pipeline, and fusing QC into a GPU aligner, with QC-speedup sensitivity rows. run_reformulation_suite.sh runs all CPU-side measurements in one go. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… to model BitstreamKind enum instead of a string, uncomp_chunk_size, RAW bitstream for Gzip/Deflate, algorithm_type sweep, median of 3 nvlzcat runs. Model gains D4: stream uncompressed output into the next tool. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- gpu_bench.py: sizes of device arrays via cupy (implicit host conversion is refused, which made every nvCOMP section record an error); free the memory pool between sections and before nvlzcat subprocesses (L4 OOM). - soa_pack.cpp: skip records whose quality line length differs from the sequence (a truncated last record read past the buffer). - reformulation_model.py: B uses the single-stream gzip rate unless --bgzf-input; picks the fastest gzip-compatible GPU compressor within 90% of libdeflate's ratio; SE BGZF pool formula (nproc - w - 3); warn without --dataset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Python API has no choice of gzip decompression algorithm, and its default decodes one single-member stream as one chunk: on the whole 2 GB file it ran for over 15 minutes. nvlzcat (lookahead algorithm) covers the full file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
They run ~1000x slower than nvlzcat on the same data (API path, not GPU capability), so they must not be picked as the GPU compressor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Option A (offload QC to a GPU) also replaces fastp's Read-object parse with an index+pack parse. D5 is that parser change alone, so A's GPU contribution can be read as A vs D5 rather than A vs base. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Survey of existing GPU/CPU work (checked against primary sources), the ways to reformulate fastp (CPU D1-D7, GPU A/B/C) with their data-transform costs, and measured results: per-read costs from the traces, transform costs from soa_pack on 4 datasets, and L4 numbers for PCIe, the QC kernel (0 mismatches vs CPU over 1M reads), single-stream and BGZF inflate, and GPU gzip ratio vs throughput. Model projections per dataset at -w 16/48 with QC-speedup sensitivity. Conclusion: the largest wins are CPU changes (fix and speed up trimBySequence's gap search, cheaper parsing, right-sized -w, uncompressed streaming); a GPU pays off only fused into a GPU pipeline, or for CPU-hours if fastp's real per-read work reaches ~10-15x per thread on the GPU. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #11. Asks whether reformulating fastp for a GPU (or other batch accelerator) pays off, counting the cost of transforming data into and out of the accelerator's format. Write-up with results:
benchmark/reformulations.md. Data:benchmark/results/reformulation-2026-09/.Tools
tools/soa_pack.cpp: CPU cost of index, pack (SoA / 2-bit) and emit, plus a representative QC kernel, on real reads.--dumpwrites a batch for the GPU check.tools/gpu_bench.py: PCIe, the same kernel on the GPU (checked read for read against the CPU), nvCOMP batched BGZF inflate, single-stream gzip (nvlzcatlookahead, nvCOMP ≥ 5.2), and GPU gzip compression.scripts/reformulation_model.py: projected CPU, GPU and wall time per reformulation from measured inputs. CPU: D1 BGZF, D2 parallel gzip reader, D3-z 1, D4 uncompressed streaming, D5 flat-batch parser. GPU: A offload QC, B GPU-resident, C fused into a GPU aligner. QC-speedup sensitivity rows included.scripts/run_reformulation_suite.sh: every CPU-side measurement in one command.Measured (n2d-highmem-48 + one L4, nvCOMP 5.3):
nvlzcat): 1.26 GB/s. That's below 16 CPU threads of rapidgzip (3.4 GB/s). The nvCOMP Python default is 31 MB/s.-z 4ratio (4.88) it runs at 0.12 GB/s, about one CPU core.Conclusion (details in
reformulations.md):trimBySequence's gap search (33% of worker CPU, plus a correctness bug), a cheaper parser (1.2–2.1× on reader-bound data), right-sized-w, and uncompressed streaming to the aligner.Self-review fixes already in:
soa_packtruncated-record over-read.Remaining
trimBySequenceplusstatReadto the GPU and measure the real speedup.measure/reformulation-build(this stack + fix: scale reader backpressure with worker thread count (#721) OpenGene/fastp#723) because upstream 1.3.7 deadlocks at-w 48.🤖 Generated with Claude Code