Skip to content

test(benchmark): full-size runs, per-stage cost model, and profiling - #9

Draft
dougnukem wants to merge 4 commits into
masterfrom
benchmark/full-size-profiling
Draft

dougnukem wants to merge 4 commits into
masterfrom
benchmark/full-size-profiling

Conversation

@dougnukem

@dougnukem dougnukem commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Checks whether subset benchmark gains hold on complete datasets, and adds a code-path model of throughput built from profiles. Built on OpenGene#725's benchmark/ directory (upstream it would stack on OpenGene#725).

Findings (details in benchmark/full-size.md)

  • Fixed stages (pre-processing, adapter detection, report) cost the same on a 4M-read subset as on the full run; processing scales with reads.
  • Projecting a subset as fixed + processing × N/n predicts full-run wall within 8% (median; naive scaling: 37%) and base-vs-head deltas within 3 points (raw subset deltas: 10). The CI benchmark in ci: thread-count correctness checks (would catch #721) and base-vs-head benchmarks OpenGene/fastp#727 now reports these projected numbers.
  • On 6 complete public runs, the stack vs upstream master at -w 16: −20% RNA NovaSeq, −17% WGBS, −5 to −8% ATAC and SE RNA, +2% miRNA (noise). Output byte-identical across builds and thread counts.
  • The largest per-read cost is adapter trimming (Matcher), 31–43% of CPU, untouched so far. Adapter detection is under 1% on full files.
  • Bottleneck model: throughput = min(reader capacity, workers / CPU-µs per read over active code paths, writer). It matches at -w 16 (150bp compute-bound, 50bp reader-bound). It fails at -w 48, where every dataset is slower than at 16 threads; that's documented as open.

Added since

  • Two-point fit beats single-size scaling. A line a + b·N through subsets of 1M and 4M pairs (first N reads of the original file) predicts full-run CPU within ~1% and wall within ~2% (median over 3 datasets × 2 core counts; worst ~7% and ~10%). Scaling one subset by N/n is off by 11% (wall) / 3% (CPU). Recompressing subsets is not the cause of the earlier overprediction.
  • 48 physical cores (Intel C4, AMD n2d, SMT off): upstream master is 1.4–1.9× slower at -w 24 than at -w 16 and hangs at 48. With the stack, nothing gains more than 6% beyond 24 workers.
  • Hardware counters (Intel C4): the workload is branch- and front-end-bound (retiring 40%, bad speculation 20%, front-end 27%, back-end 12%, memory-bound 5%). The gap-search fix with early exit runs 17% fewer instructions but 10% more cycles (branch-miss rate 2.8% → 5.2%), so an instruction-count gate would miss it. At -w 48 cycles rise only 3% while CPU seconds rise 32%, which fits an estimated ~4.0 → 3.1 GHz clock drop under all-core load.

In this PR

  • full-size.md: cost model, measured constants, two projection methods, code-path model, 48-core results, counters, runtime estimation, profiling findings.
  • results/: raw full-size, 48-core, counter and subset-scaling data.
  • Scripts: fetch_full.sh, run_full.sh + bench_full.py (timing), profile_run.sh (on-CPU/threads), profile_breakdown.py, profile_pmu.sh + pmu_breakdown.py (hardware counters).

Remaining

  • Explain the -w 48 loss beyond clock speed: off-CPU capture (BPF failed) and a pack-level trace build; AMD has no counters to check the clock-drop estimate.
  • Small-input case (tens of thousands of reads, below the detection cap).
  • Profile with merge/dedup on; counters for a short-read (50bp) dataset.
  • Explain the AMD 96-vCPU ATAC slowdown vs the earlier 24-core VM (299 s vs 213 s at -w 16): snapshot-restored disks vs 2 NUMA nodes.
  • Move CI (ci: thread-count correctness checks (would catch #721) and base-vs-head benchmarks OpenGene/fastp#727) to the two-point fit and gate on CPU per pair (done; 0.3M + 1.2M pairs, -w 4).

🤖 Generated with Claude Code

Subset benchmarks overstate speedups in stages whose cost doesn't grow with
input size (adapter detection samples at most 256K reads per mate). Adds:

- fetch_full.sh: download complete ENA runs.
- run_full.sh + bench_full.py: builds x threads x datasets on full inputs,
  with per-stage timings, CPU, peak RSS, read/base counts and an output
  digest, alternating build order and optionally from a cold page cache.
- profile_run.sh: on-CPU flame graph, perf stat, per-thread utilisation,
  off-CPU stacks and iostat for one run.
- full-size.md: per-stage cost model from the source; results, the
  subset-to-full projection check and profiling findings to follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

fastp benchmark

✅ No regressions over 10%.

base → head, median of 3 interleaved runs on one 4-CPU runner. 🔴 = worse by more than 10%, 🟢 = better by more than 5%, in both cases with no overlap between the base and head runs; unmarked changes are within runner noise.

Projected to a full-size run of 50M reads (pairs for PE). Each build runs on the first 300,000 and 1,200,000 pairs of the same input; a line through the two points gives the fixed cost and the cost per pair, and the projection is fixed + cost per pair × size. On 3 complete public runs this predicted full-run CPU within 1–2% and wall within 2–4% (median). Projected wall, CPU per pair and peak RSS are gated.

dataset -w projected wall CPU per pair (read for SE) peak RSS
synthetic_pe 4 506.07 → 504.57 s (-0.3%) 39.30 → 39.31 µs (+0.0%) 1286.23 → 1285.98 MB (-0.0%)
synthetic_se 4 246.28 → 251.40 s (+2.1%) 18.91 → 19.18 µs (+1.4%) 1240.11 → 1240.05 MB (-0.0%)
atac_hiseq_pe 4 208.92 → 207.53 s (-0.7%) 16.05 → 16.15 µs (+0.6%) 1218.70 → 1218.57 MB (-0.0%)
fixed cost and projected CPU
dataset -w fixed cost (wall) projected CPU
synthetic_pe 4 1.01 → 1.01 s (+0.7%) 1965.91 → 1966.29 s (+0.0%)
synthetic_se 4 0.66 → 0.64 s (-2.6%) 946.19 → 959.49 s (+1.4%)
atac_hiseq_pe 4 0.69 → 0.70 s (+2.2%) 803.22 → 808.26 s (+0.6%)
measured on the 1,200,000-pair subset
dataset -w wall time CPU (user+sys) adapter detection processing
synthetic_pe 4 13.09 → 13.11 s (+0.2%) 48.01 → 48.10 s (+0.2%) 0.86 → 0.83 s (-2.9%) 12.10 → 12.14 s (+0.3%)
synthetic_se 4 6.57 → 6.65 s (+1.2%) 23.44 → 23.56 s (+0.5%) 0.49 → 0.48 s (-1.9%) 6.04 → 6.09 s (+0.7%)
atac_hiseq_pe 4 5.69 → 5.71 s (+0.3%) 19.98 → 20.09 s (+0.5%) 0.52 → 0.46 s (-11.7%) 5.06 → 5.13 s (+1.4%)

- results/full-2026-09.tsv: 6 complete public runs (0.2-17 GB), upstream
  master vs the stack at -w 8/16/default and the stack at -w 48. Output is
  byte-identical across builds and thread counts on every dataset.
- full-size.md: fixed stages cost the same on a 4M-read subset as on the full
  run. Projecting fixed + processing x N/n predicts full-run wall within 8%
  (naive scaling: 37%) and base-vs-head deltas within 3 points (raw subset
  deltas: 10 points).
- A bottleneck model: throughput = min(reader capacity, workers / CPU per
  read over active code paths, writer capacity). Per-path costs come from
  profiles; it matches measurement at -w 16 (compute-bound 150bp,
  reader-bound 50bp) and fails at -w 48, which is documented.
- profile_breakdown.py: CPU-us per read per code path, saturated threads and
  implied ceilings, so the model can be refreshed after a code change.
- fetch_full.sh resumes and retries; profile_run.sh keeps off-CPU errors.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…subset fit

- results/cores-2026-09.tsv: master, stack and stack + gap-search fix at
  -w 16/24/48 on 48-physical-core Intel (C4) and AMD (n2d) VMs. Upstream
  master is 1.4-1.9x slower at -w 24 than at -w 16 and hangs at 48; with the
  stack nothing gains more than 6% beyond 24 workers.
- profile_pmu.sh + pmu_breakdown.py, results/pmu-2026-09: top-down slots,
  IPC and per-function branch misses. The workload is branch- and front-end
  bound (memory-bound 5%). The gap-search early exit runs 17% fewer
  instructions but 10% more cycles because branch mispredictions double, so
  an instruction-count gate would miss it.
- full-size.md: a line fitted through two subset sizes predicts full-run CPU
  within about 1% and wall within about 2% (median), better than scaling one
  subset. Recompression of subsets does not explain the earlier
  overprediction.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nt fit

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants