Conversation
Subset benchmarks overstate speedups in stages whose cost doesn't grow with input size (adapter detection samples at most 256K reads per mate). Adds: - fetch_full.sh: download complete ENA runs. - run_full.sh + bench_full.py: builds x threads x datasets on full inputs, with per-stage timings, CPU, peak RSS, read/base counts and an output digest, alternating build order and optionally from a cold page cache. - profile_run.sh: on-CPU flame graph, perf stat, per-thread utilisation, off-CPU stacks and iostat for one run. - full-size.md: per-stage cost model from the source; results, the subset-to-full projection check and profiling findings to follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fastp benchmark✅ No regressions over 10%.
Projected to a full-size run of 50M reads (pairs for PE). Each build runs on the first 300,000 and 1,200,000 pairs of the same input; a line through the two points gives the fixed cost and the cost per pair, and the projection is fixed + cost per pair × size. On 3 complete public runs this predicted full-run CPU within 1–2% and wall within 2–4% (median). Projected wall, CPU per pair and peak RSS are gated.
fixed cost and projected CPU
measured on the 1,200,000-pair subset
|
This was referenced Sep 29, 2026
- results/full-2026-09.tsv: 6 complete public runs (0.2-17 GB), upstream master vs the stack at -w 8/16/default and the stack at -w 48. Output is byte-identical across builds and thread counts on every dataset. - full-size.md: fixed stages cost the same on a 4M-read subset as on the full run. Projecting fixed + processing x N/n predicts full-run wall within 8% (naive scaling: 37%) and base-vs-head deltas within 3 points (raw subset deltas: 10 points). - A bottleneck model: throughput = min(reader capacity, workers / CPU per read over active code paths, writer capacity). Per-path costs come from profiles; it matches measurement at -w 16 (compute-bound 150bp, reader-bound 50bp) and fails at -w 48, which is documented. - profile_breakdown.py: CPU-us per read per code path, saturated threads and implied ceilings, so the model can be refreshed after a code change. - fetch_full.sh resumes and retries; profile_run.sh keeps off-CPU errors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…subset fit - results/cores-2026-09.tsv: master, stack and stack + gap-search fix at -w 16/24/48 on 48-physical-core Intel (C4) and AMD (n2d) VMs. Upstream master is 1.4-1.9x slower at -w 24 than at -w 16 and hangs at 48; with the stack nothing gains more than 6% beyond 24 workers. - profile_pmu.sh + pmu_breakdown.py, results/pmu-2026-09: top-down slots, IPC and per-function branch misses. The workload is branch- and front-end bound (memory-bound 5%). The gap-search early exit runs 17% fewer instructions but 10% more cycles because branch mispredictions double, so an instruction-count gate would miss it. - full-size.md: a line fitted through two subset sizes predicts full-run CPU within about 1% and wall within about 2% (median), better than scaling one subset. Recompression of subsets does not explain the earlier overprediction. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nt fit Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Checks whether subset benchmark gains hold on complete datasets, and adds a code-path model of throughput built from profiles. Built on OpenGene#725's
benchmark/directory (upstream it would stack on OpenGene#725).Findings (details in
benchmark/full-size.md)-w 16: −20% RNA NovaSeq, −17% WGBS, −5 to −8% ATAC and SE RNA, +2% miRNA (noise). Output byte-identical across builds and thread counts.Matcher), 31–43% of CPU, untouched so far. Adapter detection is under 1% on full files.-w 16(150bp compute-bound, 50bp reader-bound). It fails at-w 48, where every dataset is slower than at 16 threads; that's documented as open.Added since
a + b·Nthrough subsets of 1M and 4M pairs (first N reads of the original file) predicts full-run CPU within ~1% and wall within ~2% (median over 3 datasets × 2 core counts; worst ~7% and ~10%). Scaling one subset by N/n is off by 11% (wall) / 3% (CPU). Recompressing subsets is not the cause of the earlier overprediction.-w 24than at-w 16and hangs at 48. With the stack, nothing gains more than 6% beyond 24 workers.-w 48cycles rise only 3% while CPU seconds rise 32%, which fits an estimated ~4.0 → 3.1 GHz clock drop under all-core load.In this PR
full-size.md: cost model, measured constants, two projection methods, code-path model, 48-core results, counters, runtime estimation, profiling findings.results/: raw full-size, 48-core, counter and subset-scaling data.fetch_full.sh,run_full.sh+bench_full.py(timing),profile_run.sh(on-CPU/threads),profile_breakdown.py,profile_pmu.sh+pmu_breakdown.py(hardware counters).Remaining
-w 48loss beyond clock speed: off-CPU capture (BPF failed) and a pack-level trace build; AMD has no counters to check the clock-drop estimate.-w 16): snapshot-restored disks vs 2 NUMA nodes.-w 4).🤖 Generated with Claude Code