Conversation
Threads are named by role (fp-read-L/R, fp-work-N, fp-bgzf, fp-write, ...) so top -H, pidstat, perf and gdb attribute CPU to pipeline stages. `make TRACE=1` records timed spans per thread (read_pack, decompress, bgzf_block, process, compress, offset_wait, write, and reader/worker waits) into lock-free per-thread buffers, written as TSV at exit to $FASTP_TRACE_FILE. Without TRACE=1 the macros compile to nothing. benchmark/scripts/trace_analyze.py reports exclusive time per role, per-read CPU cost per stage, a busy/wait timeline, pack backlog and a bottleneck verdict; trace_run.sh runs it over datasets and thread counts and records trace overhead. Outputs are byte-identical to the normal build; `fastp test` passes; overhead was not measurable (median 3.82 s vs 3.83 s, 5 alternating runs). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2 tasks
- make TRACE=1 now builds ./fastp-trace from obj-trace/: -DFASTP_TRACE changes no timestamps, so switching builds in one tree silently kept the other build's objects (or mixed both). - The exit dump sets an atomic flag that stops further recording and writes a size snapshot of each buffer, so error_exit() from a worker while others still run no longer iterates vectors being appended to. - trace_analyze.py starts the window at the first reader pack and reports earlier spans separately: for BGZF input the evaluator starts its own fp-bgzf pool during detection, which stretched the window and inflated BGZF per-read cost. - trace_run.sh accepts "default" (no -w). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rch bug 16 traced cells: reader-bound in 6 of 8 gz cells, reader cost is parsing (3-5x inflate), every stage slower per read at -w 48 (24 cores x 2 SMT). perf on workers: matchWithOneInsertion 33%, trimBySequence 36% inclusive. trimBySequence's one-gap loops never offset by pos (since eb461d5, v0.26.0): fixing it trims +0.5-0.7% more reads and costs +28-42% CPU. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #10. Measures stage costs inside one run, pack by pack: which thread is the critical path, and what each stage costs per read. This is the "trace build" #9 lists as a next step.
Source changes:
fp-read-L/R,fp-read,fp-read-I,fp-work-N,fp-bgzf,fp-bgzf-io,fp-write.perf --sort comm,top -H,pidstatand gdb then attribute CPU per stage.make TRACE=1builds./fastp-tracefrom its ownobj-trace/, so normal and trace objects never mix. It uses per-thread lock-free span buffers written as TSV at exit to$FASTP_TRACE_FILE. Without it the macros compile to nothing.Analysis:
scripts/trace_analyze.py: exclusive time per role, per-read CPU per stage, busy/wait timeline, pack backlog, and a bottleneck verdict. The window starts at the first reader pack, and detection-phase spans are reported separately.scripts/trace_run.sh: runs over datasets ×-w× input kind and records trace overhead.Checked:
fastp testpasses.-std=c++11 -Wall -Wextra, with and without-DFASTP_TRACE.Results (16 traced cells;
benchmark/trace.mdandbenchmark/results/trace-2026-09/):Readobjects (670–1100 ns/read), 3–5× the ISA-L inflate.-w 48every stage is 15–40% slower per read; the serial reader shares a physical core with a worker.perfon workers:Matcher::matchWithOneInsertion33%,trimBySequence36% inclusive, compression ~20%,statRead11%.trimBySequence's one-gap loops never offsetrdatabypos(sinceeb461d5, v0.26.0), so indel-tolerant adapter matching only tests position 0. With+ pos, 2M-read runs trim 0.5–0.7% more reads, and the fix costs +28–42% CPU, so it needs a faster search alongside it. Not filed upstream yet.Self-review fixes already in: separate trace build target, exit-dump race on
error_exit(), and detection-phase BGZF-pool spans excluded from the window.Remaining
🤖 Generated with Claude Code