Conversation
Adds benchmark/: a small, reproducible corpus (two truncated public ENA accessions -- ATAC-seq SRR891268, WGS SRR952827 -- plus a synthetic generator, see benchmark/corpus.md) and scripts to sweep it across fastp releases and thread counts, so performance/correctness regressions show up in data instead of being rediscovered in production. Also backfills a real run of that sweep across every v1.3.x patch plus v1.0.1-v1.2.0 and the OpenGene#723 fix branch, on a 48-vCPU host (benchmark/results/backfill-2026-09.tsv, summarized in backfill-2026-09-summary.md). This independently confirms the root cause and fix in OpenGene#723 from timing data alone: v1.3.0/1.3.1/1.3.4-1.3.7 all show the same signature (a 45-60s stall at exactly -w 32, a hard deadlock at -w 48), v1.3.2/1.3.3 and the OpenGene#723 branch don't. Output digests agree across every version/thread/dataset combination that completed -- thread count doesn't change trimming output on either side of the fix. Version selection deliberately keeps every v1.3.x patch rather than collapsing to "latest patch per minor in the last 12 months" -- see the comment in scripts/versions.tsv for why that heuristic would hide the exact regression this backfill exists to surface.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Separate/parallel to #723 (not stacked — this is test infrastructure, independent of the fix itself).
Summary
Adds
benchmark/: a small, reproducible, license-free corpus (two truncated public ENA accessions plus a synthetic generator — seebenchmark/corpus.md) and scripts to sweep it across fastp releases and thread counts, tracking wall-clock time and an output-correctness digest per (version, dataset, threads).Also backfills a real run of that sweep — every v1.3.x patch, v1.0.1/v1.1.0/v1.2.0, and the #723 fix branch, on a 48-vCPU host — checked in as
benchmark/results/backfill-2026-09.tsv(+ a summarized.md).Why this is useful independent of #723
This wasn't wired into CI before, so the #721/#723 regression shipped, was reported by users (#705, #694), and sat for months before anyone connected it to a specific release. A cheap 1-dataset x 2-thread-count smoke sweep on every push would have caught it the day v1.3.4 shipped.
What the backfill shows
The timing data alone reproduces the root-cause story in #723 without reading a line of source: every affected version shows the same signature — a 45-60s stall at exactly
-w 32, a hard deadlock at-w 48— and v1.3.2/v1.3.3/this-fix don't:Full table (all 12 versions x 3 datasets x 4 thread counts) in
benchmark/results/backfill-2026-09-summary.md.Output correctness: zero digest mismatches across every version/dataset/thread-count combination that completed — thread count doesn't change trimming output on either side of the regression.
Version selection note
scripts/versions.tsvdeliberately keeps every v1.3.x patch rather than collapsing to "latest patch per minor in the trailing 12 months" — that heuristic would hide the exact regression this backfill exists to surface. See the comment there.Not yet done
Not wired into GitHub Actions — the full sweep (12 versions x 3 datasets x 4 threads, ~1hr on a 48-core host, several cells intentionally hang to confirm the regression) is too heavy for every PR. A 1-dataset x 2-thread-count smoke sweep against the current build would be cheap enough for CI and is a natural follow-up — happy to add it if useful.