The scripts that produced the benchmark numbers in the STAR Suite manuscript, release 1.9.5.
This repository holds the run drivers, the per-assay wrappers, the acceptance checkers, the campaign runners and the concordance analyses, as they were run, with the locations on our machines replaced by configuration variables. It is a companion to three other public repositories:
STAR-suiteis the software. The benchmarks used release 1.9.5 (tagv1.9.5, source revisionfb9f1f4e0c380d8301041cb16ff3fd9856b62ce5), built as described below; the benchmark binary's SHA256 is2ae4ed73990f81b6b44bbb8df17cd9a41adbe2c46f87aa6d5b4c0c337ee124b1.STAR-suite-provenanceholds the normalized record of each benchmark: software revision, hardware class, measured values and the caveats that bound them. It carries no host paths or command lines, which is why the scripts live here.STAR-suite-recipesholds the production workflow recipes, the same parameter contracts the benchmarks exercise.
docs/TABLE_MAP.md maps every number in the manuscript's tables to the script that produced it.
cp config/env.example.sh env.sh # fill in the locations for your machine
export STAR_PAPER_CONFIG=$PWD/env.sh
python3 helpers/build_benchmark_binary.py # build the 1.9.5 binary from the release tag
DRY=1 drivers/run_flex_arm.sh F01 # render one arm's command without running it
drivers/run_flex_arm.sh F01 # run one timed arm
python3 checkers/check_local.py F01 # accept it against the previous release
python3 campaign/run_local_campaign.py # or run the local arms one at a time, gatedEvery script reads its locations from the file named by STAR_PAPER_CONFIG. No script contains a path
from our machines. Results stay out of this repository: the drivers write to BENCH_ROOT, and the checkers
and campaign runners write their acceptance evidence to EVIDENCE_ROOT, both outside the checkout.
The drivers expect staged inputs under ${BENCH_ROOT}/stage: STAR (a copy of the benchmark binary),
STAR_upstream_2.7.11b (upstream STAR built from its tag, for the external bulk control and CellGENI),
tools/ (the release build's trimvalidate, remove_y_reads and trim_qc_fastq, its
compare_salmon_star.py and make_gene_map_from_gtf.sh, and helpers/awk_y_remove.sh from this
repository), bulk/tool_env.sh (from config/tool_env.example.sh) and bulk/tx2gene.tsv,
gene_probes_v1.1.0.tsv (cyto's gene-probe table built from the v1.1.0 probe-set CSV, included probes),
msk/ (the MSK feature references) and PBMC10K_fastqs/. Manifests record each staged file's SHA256
against its source (STAGE_MAP.tsv in LEGACY_STAGE, the index manifest read by
helpers/verify_reference_manifest.py). The drivers refuse to time an arm whose inputs are not on NVMe or
no longer match those records.
| Directory | Contents |
|---|---|
drivers/ |
one driver per assay; each takes an arm name and runs one timed benchmark under the protocol below |
wrappers/ |
the per-assay benchmark wrappers the drivers invoke (bulk, SLAM-seq, A375, MSK) |
checkers/ |
acceptance checks run after each arm, before the next one starts |
campaign/ |
campaign runners, the number-table writer, the cloud Flex runner and the GSE325982 included-probe panel |
helpers/ |
the binary build, sampling, input-audit and acceptance helpers the drivers and checkers call |
analysis/ |
the concordance scripts that produce the paper's correlation and Jaccard values, and the 320k triple-intersection analysis |
config/ |
the configuration template and the bulk tool-environment template |
docs/ |
which script produced which number (TABLE_MAP.md), and the cloud host setup (CLOUD_HOSTS.md) |
The drivers enforce the measurement protocol rather than assume it. Before a timed run each one checks the
binary by SHA256 and version, resolves every input, output, temporary and index path to its physical device
and refuses to run if any is not on NVMe, captures the RAID state and aborts on an active resync, waits for
a quiet machine (no other benchmark tool running, no process above 40% CPU, load average below 3), checks the
free space, and drops the page cache. It then runs the arm once under /usr/bin/time -v, with CPU, memory,
disk and scratch sampling. Completion is taken from the wrapped process's exit status together with a
wrapper-written completion record, not from the absence of an error. An attempt marker blocks a second
execution into the same output, and a stop marker written by any failed gate blocks every later arm.
The checkers accept an arm on parity, not bytes: its wall time must be within 1.3 times the previous release's, its executed command identical after relocating paths, and each concordance metric at least the previous release's value less 0.002 (0.2 percentage points for call agreement). Byte comparisons are recorded but do not gate, because cell calling is a Monte Carlo procedure and its marginal calls can move.
Comparator runs (Cell Ranger 9.0.1, the external bulk pipeline, cyto 0.4.7, upstream STAR 2.7.11b with the
CellGENI parameters) were timed on the same host as the STAR Suite runs they are compared with (the lab
server, or the same cloud instance type), at 32 threads, and retained across releases; docs/TABLE_MAP.md
says which run each comes from.
Local timings were measured on one server (Intel i9-13900KF, 24 cores and 32 threads, 126 GiB, Ubuntu 22.04)
with every file on local NVMe. The 320k rows and the spinning-disk comparison were measured on two AWS
instances limited to the same size; docs/CLOUD_HOSTS.md describes them. The benchmark binary was built
with gcc 11.4.0, g++ 12.3.0 and GNU Make 4.3 on Ubuntu 22.04 (helpers/build_benchmark_binary.py); the
cloud hosts ran the same binary.
The public datasets are named in the manuscript with their accessions. The input FASTQs are the original files as 10x or the archives distribute them (the European Nucleotide Archive re-compresses GSE325982's into ordinary gzip), and the CBQ inputs were converted from them once, outside the timings. The consortium datasets from the Jackson Laboratory (JAX SC2300771) and Memorial Sloan Kettering (MSK 30-KO ES) are under embargo and are not distributed; the scripts that process them are included so that their configuration is inspectable. No benchmark data, matrix, read file or result is distributed here.
MIT. See LICENSE.