Skip to content

test: replace the mainnet fixtures with the 84-block coverage set, and pack both fixture sets - #226

Open
flyq wants to merge 4 commits into
mainfrom
liquan/test/coverage-cover-fixtures
Open

flyq wants to merge 4 commits into
mainfrom
liquan/test/coverage-cover-fixtures

Conversation

@flyq

@flyq flyq commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

Replaces the 18 loose mainnet fixture blocks with the 84 blocks that coverage-replayer's set-cover (#222) selected: together they reproduce all 13,419 coverage items observed while replaying mainnet blocks 1 to 27,814,972. Both fixture sets now ship packed, as test_data/mainnet.tar.zst (5.2 MiB) and test_data/synthetic.tar.zst (239 KiB). TestFixtures parses them straight into memory once per test binary, so nothing is unpacked to disk and every existing fixture test runs over the new sets unchanged.

Contents

part files uncompressed
mainnet/blocks/: full JSON of the 84 blocks (1,277,159 to 27,784,007) 84 4.70 MiB
mainnet/blocks/: each block's parent as header-only JSON (transactions: []), the same shape as the old fixtures' parents 84 0.14 MiB
mainnet/stateless/witness/: SALT + MPT witnesses, bincode-legacy with the old fixtures' .salt envelope 168 6.22 MiB
mainnet/contracts.txt: the 622 contracts the 84 replays load, one [hash, bytecode] per line 622 lines 15.79 MiB
synthetic/: the existing synthetic set, unchanged: 22 blocks (349 to 370), 21 SALT + MPT witness pairs, 13 contracts 65 1.4 MiB

The mainnet set is 26.9 MiB raw, and 5.2 MiB (5,454,850 bytes) as tar + zstd -19; the synthetic set packs to 244,272 bytes. genesis.json is not in either archive: tests hand it to the binaries by path, so it stays in test_data/<set>/, next to mainnet's bench/, and both are unchanged. Next to it, each set gets a new manifest.txt, one <number>.<hash> line per paired block its archive must hold: the 84 selected blocks for mainnet, and the current 21 pairs for synthetic.

CI's Codespell step now skips *.zst (.github/workflows/build-and-test.yml:37). Codespell treats any file with no NUL byte in its first 1024 bytes as text, which mainnet.tar.zst happens to be, so it reported byte runs in the compressed data as misspellings.

The witnesses reference 626 code hashes. Replaying each block with bytecode loaded lazily showed that 622 of them are loaded, and all 84 blocks validate with only those 622. The other 4 belong to accounts whose balance or nonce is read but whose code never runs.

Loader

TestFixtures::mainnet_shared() / synthetic_shared() (crates/stateless-test-utils/src/fixtures.rs) read the archive entry by entry straight from the zstd stream with tar::Archive::entries() and hand each entry to the existing JSON, bincode and contracts parsers. Nothing touches disk, so there is nothing to clean up, even when a test process is killed mid-load.

  • Each set is parsed once per test binary. The owned mainnet() / synthetic() constructors and the directory loader are removed, and their remaining callers now share: the witness_encoding tests, the validator's mock RPC, a tracing-executor test and the zstd bench.
  • An entry outside the layout panics instead of being skipped. packed_sets_are_complete (crates/stateless-test-utils/src/fixtures.rs:232) requires the archive's paired blocks to equal its manifest.txt exactly, every witness to be paired, and every paired block's parent header to be loaded, so a repack cannot silently shrink or swap the fixture sweeps. Dropping the .mpt entries or the parent headers in the loader fails it, and so does removing the 27,784,007 group from the archive (reported as missing) or swapping it for another block (reported as missing plus unlisted).

New dependencies: tar (a new workspace dependency) plus zstd (already in the workspace), both in stateless-test-utils only.

Test logging

init_test_logging (crates/stateless-test-utils/src/logging.rs:13) now writes through libtest's output capture (with_test_writer()), so a passing cargo test prints no logs, while a failing test, or --nocapture, still shows them. Before, it wrote straight to stdout, which the 84-block set turned into 220 log lines per run.

Two of those were ERROR lines, keyless deploy nonce read failed … Metadata not available in witness, and they are expected. Tx 21 of block 8,064,067 is a keyless deploy of the ERC-2470 singleton factory that reverted on chain. Before Rex5, mega-evm reads the deploy signer's nonce straight from the database, bypassing the journal, and the rejected deploy merged nothing back, so the block's witness does not carry the signer. The replay's read fails into a revert with InternalError(), and since every revert on that path charges only the fixed dispatch overhead and rolls the frame back, the replay still matches the chain's receipt (status 0, 227,496 gas).

Provenance

Block numbers and hashes come from the final set-cover manifest of #222's run; test_data/mainnet/manifest.txt is that list, identical to the block table in #222's description. The data was fetched over mainnet RPC:

  • eth_getBlockByNumber for each block, with its hash checked against the manifest;
  • the header of block n−1, checked against parentHash;
  • mega_getBlockWitness for the witnesses;
  • eth_getCodeByHash for the contracts, with bytecode verified against its hash.

The one-off fetcher is not part of this PR.

Testing

  • From the packed data alone, every one of the 84 blocks passes validate_block, and also validate_block_deriving_updates anchored on its parent header's state and withdrawals roots.
  • The archive loader yields exactly the packed file counts: mainnet 168 blocks, 84 + 84 witnesses and 622 contracts; synthetic 22 blocks, 21 + 21 witnesses and 13 contracts.
  • cargo test: 492 passed.
  • cargo test -p stateless-core --no-default-features --lib: 73 passed.
  • cargo fmt --all --check, cargo clippy --workspace --all-targets --all-features and cargo sort --check are clean.
  • The witness_zstd_level bench runs, now on block 16,511,216 as the largest fixture witness.
  • cargo test -p stateless-core takes 3.2 s instead of 0.8 s with the larger set, almost all of it the executor's 84-block sweep. The witness_encoding tests drop from 0.88 s to 0.35 s now that they share one load.

🤖 Generated with Claude Code

…d pack both fixture sets

The mainnet blocks, witnesses and contracts now ship as `test_data/mainnet.tar.zst` (5.2 MiB):
the 84 blocks coverage-replayer's set-cover selected over mainnet 1..=27,814,972, each with its
parent's header, their SALT and MPT witnesses, and the 622 contracts their replays load. They
replace the 18 loose fixture blocks. The synthetic set ships packed the same way, as
`test_data/synthetic.tar.zst` (239 KiB); `genesis.json` stays unpacked in both sets, next to
mainnet's `bench/`.

`TestFixtures` parses an archive entry by entry straight from the decompressed stream, so nothing
is unpacked to disk. Entries outside the layout panic instead of being skipped, and
`packed_sets_are_complete` checks that every witness is paired and every paired block's parent
header is loaded, so a packaging slip cannot silently shrink the fixture sweeps. Both sets load
only through `mainnet_shared()` / `synthetic_shared()`, once per test binary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mega-maxwell

mega-maxwell Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Claude review status

Living comment — rewritten in place. The review workflow keeps this single comment up to date instead of posting a new one each round, so it always describes the latest reviewed commit and the earlier text is intentionally gone. No reply is needed here; reply to a finding in its own review thread, and answer an open question in a reply on this PR. The next review round reconciles your answer.

🛠️ Review did not finish

Attempted b8e99505..729367e7 · updated 2026-09-30T07:39:00+00:00

This round did not publish: MODEL_ACTION_FAILED in phase review_retry. Anything listed below is from the last round that did. Re-run the workflow or push a new commit to try again.

@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 92.7%. Comparing base (727a34b) to head (729367e).
⚠️ Report is 1 commits behind head on main.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

flyq and others added 2 commits September 30, 2026 13:57
`init_test_logging` wrote straight to stdout, which libtest does not capture, so every passing run printed the mainnet fixture sweeps' per-block debug lines, the validator integration tests' debug lines, and two expected mega-evm ERROR lines from block 8,064,067. Route it through `with_test_writer()`, so the output shows only for a failing test, or under `--nocapture`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codespell treats any file with no NUL byte in its first 1024 bytes as text. `test_data/mainnet.tar.zst` has its first NUL at byte 2187, so the lint job spell-checked its compressed bytes and failed on byte runs such as `BU` and `vEw`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread crates/stateless-test-utils/src/fixtures.rs
`packed_sets_are_complete` only checked each archive for internal consistency, so a repack that dropped a whole block group (the block, its parent header and both witnesses) or swapped one block for another still passed, and the sweeps silently ran over fewer or different blocks. Each set now keeps a `manifest.txt` of `<number>.<hash>` lines beside its archive, and the test requires the archive's paired blocks to equal it exactly. The mainnet manifest is the 84-block set-cover selection; the synthetic one lists its 21 paired blocks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants