Conversation
|
§13.11 rung 1 — quorum shadow verdict Shadow mode: this records a verdict and merges nothing. A refusal |
|
No activity since 2026-09-17 13:44Z; 2 release trains have been cut since (v0.68.2, v0.68.1). Age is the only input to this sweep — it is not a judgement on the work, and a train label such as What happens next: if one more train passes while this is still labelled To clear it: push, rebase, or say on the thread what it is waiting for. Any of the three removes the label at the next sweep. If it is blocked on something external, name that here — a blocker with an owner is not sprawl, and it stops the clock. |
|
Adopting this PR (2026-09-21 sweep; author session is not live — branch cut 09-16, no session on the box predates it). Plan, before touching anything: the DIRTY state is the squash-of-the-stack-base class — |
…annot reproduce QE2E-INV-001 could not be judged because nothing in the tree held a MEASURED Qwen3.5 tensor inventory to judge against. This adds one: the 320 tensors of ~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782...dc517), read straight from the GGUF header rather than from a model card. It already falsifies the current arithmetic. Dense/GQA accounting applied to that file gives 644,400,128 against a measured 752,393,024 — short by 107,992,896, 14.4% of the model, because 18 of the 24 layers are Gated DeltaNet and no term here counts their conv, gate, state or output projections. Two shapes in the file are also not what dense accounting predicts, and both are pinned: attn_q is [1024, 4096] = 2 * num_heads * head_dim (the q projection emits the attention output gate alongside the query; attn_output [2048, 1024] confirms num_heads * head_dim = 2048), and attn_q_norm/attn_k_norm are present at head_dim. The file is also TIED — it has no output.weight. Refs #3346 Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… a hybrid model can be counted contracts/model-families/qwen3_5.yaml declares inner_size, state_size, conv_kernel, group_count and full_attention_interval under constraints:, and ModelConstraints carried none of them. A Gated DeltaNet layer's parameters live entirely in those dimensions, so every consumer of the descriptor counted Qwen3.5 as if three quarters of its layers did not exist. Carried through as ModelConstraints::deltanet: Option<DeltaNetShape> — the runtime YAML loader (parsing.rs) and the compiled-in registry (build_parsing.rs + build_codegen.rs) both populate it, and FALSIFY-MF-QWEN35-010 pins that the declared values survive the trip and that no other family acquires a shape it never declared. qwen3_5.yaml is the only descriptor with these keys, so every other family keeps byte-identical accounting. model_arithmetic gains gated_deltanet_layer_params (one term per GGUF tensor: attn_qkv, attn_gate, ssm_conv1d, ssm_alpha/beta, ssm_a, ssm_dt.bias, ssm_norm, ssm_out) and hybrid_layers (the interleaved schedule). attention_layer_params gained two terms the real file has and dense accounting did not model: the gated q projection (2*n_h*d_k) and the q/k norm vectors. Falsified against a real model, not against itself: fed the 0.8B configuration, the equation now reproduces the 320-tensor inventory of Qwen3.5-0.8B-Q4_K_M.gguf EXACTLY — 752,393,024, both layer kinds matching tensor for tensor. QE2E-INV-001 is still NOT asserted, and no range was widened to make it pass. The 9b descriptor now yields 8,344,907,136, up from 8,208,519,168 but still 0.655B below [9.0B, 9.2B]. The remaining gap looks like descriptor drift rather than missing arithmetic: 9b keeps inner_size 2048 — the value the 0.8B uses at hidden_dim 1024 — while quadrupling hidden_dim, and its group_count 8 fails 8 * 128 == 2048, a consistency the measured 0.8B satisfies at 16 * 128. Only a real Qwen3.5-9B file can settle it; none is on this box. Refs #3346, #3347 Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…osed The note said the obligation was undischarged because ModelConstraints does not carry the DeltaNet shape keys. It does now, and the 0.8B count reproduces the real GGUF exactly. What actually blocks the obligation is descriptor drift at the 9b variant. Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd pv extract shrank the graph by 356 triples instead of refusing d3cc76f rewrote the note on the QE2E-INV-001 binding and dropped the trailing `"` — contracts/binding.yaml stopped being valid YAML at line 764 (`found unexpected end of stream`). Nothing in the PR noticed because `pv extract contracts` does not refuse a binding registry that will not parse: it emitted a graph with 15,244 triples where main has 15,600 — every bound symbol AFTER the broken entry (prune::run, distill::run, harness_ir::*, ptx_explain::run, …) silently gone — and `--check` would have agreed with itself. Found while regenerating the derivative for this adoption, by the drop, not by any gate. One character. With it, binding.yaml parses (156 entries, same as main) and the extraction is byte-identical to main's committed contracts.nt, so this PR owes no graph change after all. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
d3cc76f to
da1ee15
Compare
|
Adopted and rebased — and the rebase found a real defect in the branch, now fixed. What was done. What the rebase found. Regenerating Verified on the new head: Not armed: per the board ruling every PR needs a quorum receipt first; that runs next. Refs #3346. |
…t to judge against Acceptance transcribed from issue #3346 as this PR answers it (the type that reads the descriptor was wrong, not the range or the descriptor), with the 9B range instantiation explicitly out of scope until a real 9B GGUF exists. Refs #3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
quorum-review (AD-04): NOT agreed (auto_merge: checked=true was_armed=false disarmed=false) {
"ticket": "PMAT-3346",
"head": "1ffdccf2500116a4dd39af6a66a697f7bb841136",
"width": 3,
"executor": "agy",
"agreed": false,
"auto_merge": {
"checked": true,
"was_armed": false,
"disarmed": false,
"note": "auto-merge not armed"
},
"lanes": [
{
"lane": 1,
"verdict": "FAIL",
"findings": 2
},
{
"lane": 2,
"verdict": "NO-VERDICT",
"findings": 0
},
{
"lane": 3,
"verdict": "PASS",
"findings": 4
}
]
} |
…iases Found by the AD-04 quorum on #3350 (lane 1, gemini-3.1-pro-high, cited model_arithmetic.rs:144): the projection term used q_out for a gated family's q matrix (2*n_h*d_k — attn_q emits the output gate, MEASURED in Qwen3.5-0.8B) while the bias term still used q_dim. A bias narrower than its projection is not a model. One token: q_dim -> q_out in the bias sum. Why a delta-0 measurement did not catch it: no shipped family exercises the case. Qwen3.5 has no attention bias; Qwen2.5 has biases but is not gated, so q_out == q_dim there. The test that pinned 12 (a q_dim bias under a q_out matrix) now asserts 16 and says why, and a second test holds the other polarity — a non-gated family with biases is unchanged at 12. 115 model_arithmetic + model_family tests pass; oracle 218; clippy clean. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Quorum round 0: 1 FAIL / 1 no-verdict / 1 PASS — and the FAIL was a real finding, now fixed in |
|
quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false) {
"ticket": "PMAT-3346",
"head": "e060c7e2c22bc221b5e454a9b677d09d2b27c2e2",
"width": 3,
"executor": "agy",
"agreed": true,
"auto_merge": {
"checked": true,
"was_armed": false,
"disarmed": false,
"note": "auto-merge not armed"
},
"lanes": [
{
"lane": 1,
"verdict": "PASS",
"findings": 5
},
{
"lane": 2,
"verdict": "PASS",
"findings": 0
},
{
"lane": 3,
"verdict": "PASS",
"findings": 2
}
]
} |
Cross-inspection of
|
| lane | conversation | status | verdict | findings | refs to sibling lanes / receipts / $WORK | duration |
|---|---|---|---|---|---|---|
| lane 1 | f3ea1f1c |
SUCCESS | PASS (structured_output) | 5 (2 cited) | 0 | 644 s |
| lane 2 | 3f01ebe3 |
SUCCESS | PASS (structured_output) | 0 (0 cited) | 0 | 55 s |
| lane 3 | 81760200 |
SUCCESS | PASS (structured_output) | 2 (0 cited) | 0 | 394 s |
Distinct agy conversation ids: 3/3; no lane references a sibling lane, another receipt or $WORK. Read from /tmp/claude-1000/-home-noah-src-aprender/51506a7a-265a-4e5c-bc53-50274de4a757/scratchpad/lanes-3346-round1 on this box. Note: round 0's pro lane found the gated-family bias-term defect (model_arithmetic.rs:144), fixed in e060c7e; this is round 1 on the new head; d4 armed via pmat-merge on the parent rule and this inspection is confirmatory.
Verdict line: 3/3 PASS, independent. Arm confirmed.
|
Folded into the 0.69 release batch #3669 by the cop (08:05Z, operator: "most PRs can be batched"). One CI run and one queue slot for all of them, and the generated files ( |
|
Landed in #3669 (squash |
…run, one queue slot (paiml#3669) * fix(pv): --table panicked on the real corpus — it cut a property by BYTE index (#3338) `pv proof-status contracts/ --table` exits 101 on `contracts/`: byte index 40 is not a char boundary; it is inside '∈' (bytes 39..42) of `for Q4_K_M Qwen2.5-Coder, quantization ∈ {Q4_K, Q6_K}` The column width is a byte count (`property.len()`, capped at 40) and `truncate` sliced `&s[..max]`, so any property whose byte 40 lands inside a multi-byte char panics. Eight contracts in `contracts/` do; the first one walked is `apr-inspect-quantization-v1.yaml`. The budget stays a byte budget — the table is laid out in bytes — and the cut now walks back to the nearest char boundary. Two tests, both RED before this commit (each panicked at obligation_matrix.rs:169): the helper row, and one through `format_obligation_table` because that is the path the operator hit. The helper fixture asserts `!s.is_char_boundary(40)` first, so it cannot silently stop proving anything. Pmat-Ticket: PMAT-3347 Refs #3347, #3338 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pv): the L2 column read an INDEX, not a link — so 3,573 of 3,753 obligations ticked (#3347) `obligation_matrix` computed `l2_tested` as `idx < falsification_tests.len()`. Obligation 3 was "tested" because the contract happened to hold 4 tests, whoever those tests were about. All 7 obligations of `qwen35-e2e-verification-v1` showed ✓ before a single test existed. The fallback was a substring match between the obligation's `property` prose and a test's `rule` prose, which is an inference, not a claim. WHICH LINK, decided by counting the corpus (1,842 files, 3,792 obligations, 4,691 falsification tests), not by preference: proof_obligations[].discharged_by 89 (62 `falsification_tests[N]`, 13 a test id, 1 a YAML sequence, rest comma lists / prose / a kani id) falsification_tests[].binds_to 38 falsification_tests[].obligation 26 (12 name an obligation id, 6 the exact property text, 8 dangle) proof_obligations[].id 180 of 3,792 kani_harnesses[].obligation 1,918 — that is the L3 column, not this one All three L2 spellings are DECLARATIONS by the contract author, so all three are read. `binds_to` is a serde alias of `obligation`, which is safe only because no entry carries both keys — checked, because serde turns that into a `duplicate field` parse error rather than a silent pick. Not read: `applies_to`. 12 `binds_to` values match one, but `AppliesTo` is an enum with `#[serde(other)] Other`, so the string is discarded at parse and there is nothing left to compare. Those 12 report `?`, not a false ✗. WHAT THE COLUMN NOW SAYS. Three values, because "no test covers this" and "nothing here says which test covers what" are different facts: ✓ Tested a test in this contract cites this obligation ✗ Untested the contract's links resolve and none names it — or it ships no falsification test at all, which is a reading, not a gap ? Unknown no readable link; not measured. An unread window is Unknown, never a tick MEASURED over `contracts/`, and the drop IS the point — it is what the old column was hiding: before after L2 ✓ 3,573 86 L2 ✗ 180 65 L2 ? 0 3,602 (3,753 obligation rows, 873 contracts) Nothing was adjusted to keep the number up, and no threshold was added. The two link fields did not exist on the structs, so both keys were written to disk and silently dropped on parse — the shape of #3314 (`id`) and #2465 (`test_harness`). `discharged_by` is typed `Citation` (scalar | comma list | YAML sequence) because `Option<String>` failed the WHOLE corpus on `publish-manifest-v1`: `invalid type: sequence, expected a string`. RED first: `l2_does_not_tick_for_an_obligation_no_test_cites` — two obligations, two tests, both citing OB-A. It asserted through the rendered table (the surface that was lying, and API-stable), so it is the SAME test before and after: on the old code `Beta holds | ✓`. Its OB-A arm keeps a fix that merely stopped ticking everything from passing. Out of scope_paths, and forced: `lint/strict_test_binding.rs` holds the only EXHAUSTIVE `FalsificationTest` literal in the tree (no `..Default::default()`), so no schema field can compile without that one line. Pmat-Ticket: PMAT-3347 Refs #3347, #3091, #3114 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pv): single-file --strict-test-binding reported every ref missing — it now refuses (#3347) The gate resolves cited test names against a source index rooted at the contract path's PARENT. The directory form gets the repo root and finds `crates/`; the single-file form gets `contracts/`, which holds no source, so every cited ref resolves to nothing and all of them are reported missing. Measured on `contracts/pv-artifact-kinds-v1.yaml`, a control whose eight refs all resolve: pv lint contracts/pv-artifact-kinds-v1.yaml --strict-test-binding total_refs 8, existing 0, missing 8 pv lint contracts/ --strict-test-binding total_refs 548, existing 521, missing 27 <- none of the 27 is this one A control contract failing identically to a broken one is a gate that cannot discriminate, and it fails SILENTLY: `passed` is true in non-strict mode, so the run still says `Result: PASS` while printing eight false findings. REFUSED rather than repaired. The scan root is computed in `provable_contracts::lint::run_lint`, outside this ticket's scope; a refusal lives in the caller, is honest, and cannot be mistaken for a clean bill. If the root is later made explicit (a `LintConfig` field the CLI can feed, which is what `--crate-dir` does NOT do today), this refusal is what should be deleted. Exit 1, not the exit-2 `decline:` class: exit 2 belongs to `ZeroContracts`, whose message ("0 contracts under ...") would be false here — there IS a contract; it is the gate that cannot run over it. Two tests, in the CI-wired `cli_integration` target: the refusal names the flag and prints no PASS and no findings, and a directory-form control proves the refusal is specific to the single-FILE form rather than to the flag. Pmat-Ticket: PMAT-3347 Refs #3347 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(crux): AutoGluon becomes a CRUX competitor — category O, 24 contracts, 25 tickets, registry edit + mutation proof Research of ../autogluon (1.6.3 @ 77946149) as a competitive-research source for aprender, landed the way category N landed linfa and burn (#3169): the competitor is admitted to the CLOSED registry, every story is a real contract with falsification gates, every story is a registry row, and every missing/partial story has a GitHub issue and a roadmap fragment. What AutoGluon is, measured from the tree rather than recalled: three predictors. TabularPredictor (65 public methods, 11 presets, 24 model families of which 8 are tabular foundation models added in 1.4-1.6), TimeSeriesPredictor (Chronos-2/Toto-2 pretrained, 30+ local/deep models, 16 metrics incl. WQL/MASE/RMSSE, auto backtesting since 1.5) and MultiModalPredictor. Evidence under evidence/crux/autogluon/. What aprender has, measured at eb262f8eb: automl/ is a single-estimator hyperparameter tuner (TPE, grid, random, DE, TimeBudget, EarlyStopping); time_series/ is one univariate ARIMA; encoders, calibration, SHAP/LIME/ permutation importance and KFold/cross_validate exist as building blocks. No predictor-level fit(label), no leaderboard, no bagging, stacking or greedy weighted-ensemble selection, no panel forecasting, no quantile forecast metrics. The gap is the AutoML UX, not the algorithms. Category O — AutoML Parity — 24 stories: 9 P0 (the README hello-world: fit(label), problem-type inference, presets, leaderboard, feature pipeline, weighted ensemble, time budget, panel forecaster, quantile metrics), 9 P1 (bagging, stacking, importance, threshold calibration, deployment artifact, tabular foundation model, backtesting, local baselines, pretrained forecaster), 6 P2 (refit_full, distill, infer_limit, fit diagnostics, memory-aware fit, covariates). MultiModalPredictor, autogluon.cloud, MLZero and Ray-parallel fits are CUT on the epic with reasons. CRUX_COMPETITORS: [&str; 14] -> [&str; 15] + autogluon Not a BEAT pillar: aprender claims no pinned-benchmark win over AutoGluon. Both registry tests that keep BEAT_INCUMBENTS and CRUX_COMPETITORS apart are extended, not worked around. Tickets: epic #3370, stories #3371-#3394, label pareto-autogluon. Roadmap: 25 fragments under docs/roadmaps/entries/, roadmap.yaml regenerated by the aggregator (idempotent check passes). Spec: docs/specifications/crux-competitive-research-ux-workflows.md v2.2 -> v2.3 — §3 gains rows for linfa+burn (category N, which #3169 never recorded there) and AutoGluon; §5 gains Category O; §6 notes that coverage_intake in the YAML is the source of truth. coverage_intake 267 -> 291 (partial 72 -> 77, missing 152 -> 171). Verification: - pv built from THIS tree validates 25/25 (24 new + master). The stale ~/.cargo/bin/pv rejects crux-O-01 with CRUX-002 — the behavioural delta proves the registry edit engaged. - Mutation-verified: deleting "autogluon" from CRUX_COMPETITORS turns competitor_registry_covers_the_corpus_vocabulary RED with "autogluon is used by contracts/ and must stay in CRUX_COMPETITORS" and the_real_crux_registry_rows_are_all_in_domain RED. Restored: 20/20. - cargo test -p aprender-contracts --lib: 1526 passed, 0 failed. - Every falsification gate is LIVE-PENDING prose (no `::`), so strict-test-binding has nothing to refuse; the obligations are RECORDED as unfalsifiable-by-absence, not satisfied. - README CONTRACT_COUNT regenerated 1835 -> 1866 by readme_sync.sh. - Guards: roadmap fragment/ids/sorted/additive/completion, contract test-binding and enforcement, shell-lint ratchet, hardcoded paths, readme claims, grep -q ratchet — all rc=0. Pmat-Ticket: PMAT-3370 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(crux): bind the admission PR to its own ticket PMAT-3401 — the epic's acceptance criteria describe the programme, not this diff Round 2 of the quorum read the epic (PMAT-3370) as the ticket and refused the admission for not implementing the 24 stories. The admission is its own unit of work with its own done-when; this fragment says so. Closes #3401 Pmat-Ticket: PMAT-3401 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(crux): PMAT-3401 title inventories the diff exactly — 26 fragments (its own included) and the scaffold CATEGORY_NAMES edit Quorum round 3 lane 1 refused on two literal mismatches between the ticket and the diff: '25 roadmap fragments' (there are 26 once this ticket's own fragment lands) and an unlisted edit to scripts/crux_scaffold_contracts.py. The title now lists every path. Pmat-Ticket: PMAT-3401 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * roadmap(PMAT-3495): VERIFY-001 on 0.71.0 — run the 132 Kani harnesses in CI (proof credit from runs, not declarations), then a Verus pilot on one dequant/parser function (#3495) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * quorum(PMAT-3496 brief, PR #3496): 3/3 PASS on 57e0e4dfb — gemini lanes, measured; author claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * PMAT-3577: the logit-parity receipts under contract — parity-receipt-v2, extract:parity-receipt, and the 7 back-filled Until this row the logit-parity records under evidence/parity/** had NO validator of any kind. Not a weak one — none. That is why seven of them sat in the tree carrying no comparator for months: there was nothing that could have noticed. WHAT LANDS · contracts/parity-receipt-v2.yaml — three shapes: parity-receipt-complete (closed, ignoredProperties empty), parity-comparator-self, -oracle. The subset has no sh:or, so the comparator split is two shapes over two subclasses the extractor assigns by kind. · contracts/parity-receipt-v1.yaml — the retired layout, recorded with NO shape: all instances were migrated, and a shape whose target class nothing instantiates passes vacuously. The enforcement that replaces it fires — the extractor refuses an unmigrated record BY NAME, and so does the predicate. · ontology/extract/parity_receipt.rs — record / unmigrated / other, and skipping is never silent. · scripts/parity_receipt_denominator.sh + evidence/parity/EXPECTED_RECEIPTS — the count is pinned by an INDEPENDENT predicate. An extractor checked against a number the extractor produced proves nothing. · Unknown{ExtractorMiss}, exit 2 — a new element of the verdict lattice. An extractor that silently saw the wrong corpus reports the same "no violations" as one that saw all of it. · the 7 records migrated to v2 and back-filled with comparator {kind: self, reason}, in this commit, as the row requires. THE DENOMINATOR IS 7, NOT 8. The #3574 receipt is in PR #3575, still open; it is not on main. Measured: 113 files under evidence/parity, 7 parity records, 0 with a comparator. #3575 bumps it to 8 when it lands — this row's own falsifier on its first real use. check_parity_receipt.sh IS NOT TOUCHED, and that is the finding, not an omission. It validates the THROUGHPUT family (instrument, protocol_ref, lanes[], decode_tok_per_sec, the #2696 cross-class defect); a logit record has never carried one of those keys. Folding it in — as item 6 asked — would have deleted the #2696 validator from a family nobody was watching. One validator per artifact family, and the discriminator is the artifact's required keys, never its filename. THE BACK-FILL IS A RELABEL. `raw` is the original apr parity --json document key for key; every envelope field is quoted from committed evidence named in each record's provenance.record. model_sha256 is carried only by the one record that measured it: hashing the files today and attaching that to a receipt about 2026-09-06 would be a claim about a different world wearing a witness's clothes. One derived field was wrong first time and is worth recording: result.verdict copied raw.parity — apr's own per-position flag — which says PASS for the two 1.5B cells their own RECORD.md calls RED. It now resolves the threshold from thresholds.yaml and reproduces all seven readings the RECORD.md files state, both REDs included. No threshold is ever typed into a shape. NOTHING IS ARMED. armed_shapes lives in lint-baseline.json, a shared file this row may not touch (decision 7); the shapes are computed and reported, as ladder-green was at ONT-4c1. Arming is a follow-up with the label. CONTROLS, both directions: 7 violations with the comparator stripped from all seven → 0 as committed; exactly 1 for a single plant, naming focus node and property; widening sh:in to accept `oracle` turns ont4c3_parity_receipts RED, and mutating only the real contract turns the fixture-drift test RED; an unmigrated record declines at exit 2 naming the file; 2 receipts against a denominator of 1 declines naming both numbers; pass / fail / decline are 0 / 1 / 2. cargo test -p aprender-contracts --lib ontology::extract::parity_receipt 10 ok cargo test -p aprender-contracts-cli --test ont4c3_parity_receipts 8 ok bash scripts/parity_receipt_denominator.sh --self-test 4 ok pv lint contracts --gate {sigma,relations,shapes} Pass, 0 violations pv extract contracts --check rc 0 ONT-4c3 is NOT bound in the ONT-001 ledger: that ledger is paiml/infra's paiml-ontology.md, where v4.8 still defines ONT-4c3 as KERNEL receipts. The re-scope is an infra PR, not this one — raised with the cop rather than left as a checked box. Receipt: docs/audits/impl-PMAT-3577-receipt.md. Refs #3577, #3576, #3269, #3575, #3567 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: name it WrongCorpus, not ExtractorMiss — a prefix of its opposite is a defect `ExtractorMissing` already meant the opposite thing: an extractor that does not exist. `ExtractorMiss` would have sat beside it in the same 18-element lattice, one letter apart, with the shorter a PREFIX of the longer — `grep ExtractorMiss` matches both, and any substring test over the reasons merges them silently. That is the defect class this repo keeps paying for (#3573 today), so the name now says what happened: the extractor RAN and read a corpus the tree does not declare. Renamed before it reached main, when a rename is a sed and not a migration. Caught in review by the cop. Refs #3577, #3573 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: the ONT-4c3 divergence is filed as paiml/infra#814, not left in a thread Two repos holding different definitions of one identifier is a row, not a flag in a message: nothing collides where a tool would see it, so it collides in a person's head months later when they implement the ledger's meaning and find their correct work unusable. Refs #3577, paiml/infra#814 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * #3605: removed_by gets a shape — refusal-receipt-v1, with a closed sentinel set so the required field manufactures nothing `removed_by` had 0 occurrences tree-wide, no schema and no validator. Every refusal #3597 writes would have minted an unenforced convention, and a tree full of consistent-looking `removed_by:` lines reads as validated when it is decoration. PHASE 0 — ITS OWN CONTRACT, AND NOT BECAUSE IT IS EASIER TO WRITE Option B is refused on a fact: `parity-receipt-v1` DOES NOT EXIST AT HEAD. It is in unmerged #3600, and there it is deliberately the RETIRED layout with no shape and no instances. A live validated field on a superseded contract with zero focus nodes is a field nothing can carry. Independently, by the discriminator this tree has now used twice — an artifact family is named by its REQUIRED KEYS, never by its filename — a parity receipt requires host, backend, comparator, threshold_source and per-position metrics; a refusal requires a verb, a reason, an exit code and removed_by. They share no required key. A THIRD option was considered and refused, and it is the more attractive one: put removed_by on apr-cli-commands-v1.yaml, where the verb already IS the focus node and the universe is already the trustworthy 111. Refused because the registry is the UNIVERSE and #3597's method is to DIFF the registry against the buckets. If the buckets live in the registry, the denominator and the numerator are the same artifact and the diff is vacuous by construction. The registry says what EXISTS; a refusal says what was TRIED. Keeping them apart is what lets that diff be a real diff. THE FORCED-BINDING TRAP, AND WHY minCount 1 IS SAFE HERE A required field with no escape manufactures false data: `pv validate` requires kani_harnesses so authors fabricate one, and apex#57 had lean_theorem copied verbatim into twelve contracts resolving to nothing, gate green throughout. The escape is a closed sentinel set, so "there is legitimately nothing here" is SAYABLE and still CHECKABLE: v<major>.<minor> the release that removes it never refused permanently by design unscheduled a defect with no release chosen tbd, soon, n/a, pending, a bare `0.70`, a patch-level `v0.70.1` and a git sha are all RED. A version rather than a sha because a refusal answers "which release do I need?" and a sha is precise about the tree and silent about the boundary; the sha form is made INVALID rather than discouraged, because a shape that permits two spellings gets both. BOTH ARMS, and the second is the point: 9 cases over 8 fixtures. refusal-ok carries all three declared forms and passes with 3 focus nodes; refusal-undeclared-sentinel plants the PLAUSIBLE `tbd` and must go red. Mutation proved red-capable: widening the pattern to accept tbd fails exactly a_plausible_but_undeclared_sentinel_is_refused and accept_and_refuse_are_distinct_answers, and restoring makes 9/9 green. Fixtures carry the real contract byte for byte, so widening it without them is caught too. No new extractor: entity {type: json, ref} + vocabulary is the existing machinery for "validate this document", and a second reader for one more family is the thing this tree keeps filing against. done_when 6 — the interim re-check list is EMPTY: `removed_by` still has 0 occurrences at HEAD, so no refusal written under the interim needs revisiting. The ledger is seeded with the one refusal already measured (`apr bench` refuses qwen35) so the shape has a real focus node instead of passing vacuously over an empty list. WHICH verbs land there is #3597's bucket, not this row's. Refs #3605, #3597, #3600, #3080 Pmat-Ticket: PMAT-3605 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: count parity-receipt in by_entity_type — ONT-001 v4.10's probe reads ABSENT otherwise The spec's probe asks `by_entity_type["parity-receipt"] == 7`. Measured on this branch before the change: the map carried pv-contract, gguf, apr-model, code and lean, and NO parity-receipt key. The seven records were there — by_shape showed parity-receipt-complete=7 — but the entity-type map did not carry them, so the probe would have read ABSENT. AN ABSENT KEY IS NOT ZERO. A consumer treating it as one measures nothing and calls it a pass — the same shape as #3610, one map over. Registering the entity type in Sigma was not enough; it has to be counted where the probe looks. A test now asserts the key exists AND that it equals the shape's own focus-node count, so the two numbers cannot drift apart. Refs #3577, #3610, paiml/infra#831 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: regenerate contracts.nt with a pv built from THIS worktree — the shared target dir served another branch's binary The merge commit's contracts.nt was 107 triples short: every parity-receipt and parity-comparator node was missing, because `cargo metadata`'s target_directory is shared across worktrees and the pv on PATH had been built from a different branch minutes earlier. AND `pv extract contracts --check` PASSED ON IT, because the check re-derives the graph with the same binary. A stale tool comparing an artifact against its own re-derivation agrees with itself about nothing being there — the derived file and the checker were wrong in the same direction, which is the only way that gate can fail to fire. Rebuilt with CARGO_TARGET_DIR pinned to this worktree; the 107 triples return and by_entity_type[parity-receipt] reads 7 rather than absent. Refs #3577 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * register the new CLI test in scripts/tree_reader_tests.txt The registry is derived from the tree and drifts the moment a test that reads the tree is added without listing it. Each of the three new tests failed this guard on its OWN branch — not a shared commit, as first read. Pmat-Ticket: PMAT-3598 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * register the new CLI test in scripts/tree_reader_tests.txt The registry is derived from the tree and drifts the moment a test that reads the tree is added without listing it. Each of the three new tests failed this guard on its OWN branch — not a shared commit, as first read. Pmat-Ticket: PMAT-3598 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * #3600: the migration broke TWO readers of the records, including the release judge I asked "who else reads this?" for the <unk> assumption and did not ask it for my own migration. Two consumers read the legacy top-level metrics: crates/apr-cli/src/commands/parity_admission.rs -> mac-check RED scripts/check_model_parity.sh -> C14, the RELEASE judge The second is the one that matters: C14 is the pre-publish dogfood's parity gate, and on the migrated records it reported "no per-position metrics in the output" for a record that is fine. The migration would have taken the release gate down. THE RULE, STATED ONCE IN BOTH READERS: a v2 receipt EMBEDS the raw `apr parity --json` document under `raw`, so look inside the envelope when there is one. A fresh `apr parity` run is the raw document itself and carries the readings at the top level. Those are two different INPUTS — a tool's output and an archived receipt quoting it — not two spellings of one, which is the distinction that makes this a rule rather than the permissiveness #3613 refuses. The self-test's fixture builder read the same way, so the must-RED twin was being built from a KeyError and that control could not have fired. Verified: 7B sentinel PASS, 1.5B sentinel RED (unchanged from pre-migration), check_model_parity.sh --self-test 26/26, parity_admission 19 tests. Refs #3600, #3577 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3604: the F2 hybrid guard runs once per (model sha256, apr version, device) and leaves a receipt f2_validate_qwen35 proves the CUDA hybrid forward against a 64-position CPU reference before it will serve a token, on EVERY apr run. Measured (#3598 row 1): 67 % of a 14 s time-to-first-token, 90-93 % of it the CPU forward. The guard is right to exist and wrong to run per call. Now: the first run of a (model sha256, apr version, device) triple validates and writes a receipt; a later run whose triple matches reads it and skips the forward; `apr run --revalidate` forces a fresh run and rewrites it. MEASURED HERE, RTX 4090, Qwen3.5-4B-Q4_K_M, the 144-word row, --max-tokens 1, GPU occupancy recorded before every run, binary built from this tree under a PINNED target dir (the shared one handed me another worktree's binary first): --revalidate, warm cache 17.18 s wall guard 9,747 ms on 65 positions receipt hit, warm cache 7.39 s wall guard 0 ms sha256 1,173 ms -9.8 s wall; guard 9,747 -> 0; the key costs 1.17 s/run (2.74 GB at 2.3 GB/s) and is printed separately so it cannot hide in either number. THE RECEIPT IS THE VALIDATION, WHICH IS WHY IT IS STRICT. Every path that is not "three keys match" validates, and the three planted-receipt falsifiers were run END TO END in the real binary, not only as unit tests: wrong model sha256 -> re-validated: "receipt is for model bbbb…, this file is 00fe…" wrong apr version -> re-validated: "receipt written by apr 0.61.0, this is 0.68.2" wrong device -> re-validated: "receipt written for NVIDIA GB10, this device is …4090" corrupt file -> re-validated: "receipt unreadable (…: not a receipt)" missing file -> re-validated: "no receipt for this model" Absence is never consent, and absence and unreadability are told apart. A RECEIPT IS WRITTEN ONLY AFTER A VALIDATION THAT JUDGED SOMETHING. The guard has three early exits that let the GPU serve without comparing a position — SKIP_PARITY_GATE=1, a probe under two tokens, a CPU reference that would not run. f2_validate_qwen35 now returns F2Verdict {Accepted{positions_judged}, Rejected, NotJudged} instead of bool, and only Accepted writes; otherwise a one-token prompt would "validate" the triple for every prompt after it. The decision table (f2_receipt.rs) is pure and CUDA-free, so its 13 tests run on every build. --revalidate reaches the guard through the same env seam the guard already reads SKIP_PARITY_GATE from, rather than a 38th positional parameter on run_entry::run and six forward signatures #3606 is changing. done_when 5 is partial and says so: [source=receipt|fresh] is on the guard's stderr line; the `apr run --json` field lands with #3606's StageTimings, and F2Outcome{source, validate_ms, sha256_ms, receipt_path} is returned to the call site for exactly that. #3606 and this PR both edit f2_validate_qwen35's return path; whichever lands second reconciles ~10 lines, and #3606's lane is told. Evidence: evidence/perf/3604/MEASUREMENT.md + the stderr of all nine runs. Refs #3604, #3596, #3598, #3606 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * roadmap: mint PMAT-3604 so the AD-04 quorum for #3634 can run pmat work status PMAT-3604 refused ('Item not found'): #3604 was minted as a GitHub issue with a done_when but never as a roadmap entry, and the quorum script hard-requires the work item. Fragment + aggregate, nothing else. Refs #3604, #3634 ont-delta: none — a roadmap entry; no ontology entity, shape, reason or resolution Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(run): --gpu that fell back to CPU reported success — the refusal had no caller (#3602) `apr run --gpu` on a model whose GPU attempt is rejected at runtime printed a result and exited 0. Measured on lambda (RTX 4090 sm_89, apr 0.68.2 e6f77c98c, qwen2.5-coder-0.5b-instruct-q4_k_m): exit=0 stderr: warning: GPU output diverges from CPU at position 1 (cosine 0.4153) stdout: { "used_gpu": false, "inference_time_ms": 33646.13 } 33.6 seconds on the CPU, reported as a successful --gpu run, with `used_gpu: false` as the only signal — the same value a deliberate CPU run reports. The decision was already made, recorded and unit-tested. `registry::after_generation` implements it (R-0b, #3002/#3042): a FORCED accelerator that fell to CPU is a refusal, a DEFAULT selection that fell to CPU gets a corrective line. A git grep found its only callers were its own tests. `registry::announce` and `registry::parity_line` are in the same state, which is why a real run prints zero `selected:` and zero `parity:` lines. So this is not a new policy. It is the recorded one, reached from `apr run`: - `dispatch.rs` classifies the request ONCE via `registry::Request::wanted()` rather than re-deriving "forced" — two spellings of one rule drift apart. - `run_entry::run` calls `reconcile_accelerator` before any success output. - `--json` gains a `backend` object: `requested` / `ran` / `fell_back`, because `used_gpu: false` alone collapses "ran on CPU deliberately" with "asked for the GPU and was refused it". `accel::tests::every_accelerator_surface_calls_the_refusal` passed throughout: it guards `ensure_available`, the BUILD-time refusal, not the RUNTIME one. A guard over one of two refusals reads as coverage of both. Falsifier, both directions (`run_tests_accel_reconcile.rs`, 6 cases). Planting the pre-fix behaviour (`reconcile_accelerator` → `Ok(None)`) turns `a_forced_accelerator_that_ran_on_cpu_is_refused` and `the_json_distinguishes_a_deliberate_cpu_run_from_a_rejected_gpu_run` RED while `a_forced_accelerator_that_actually_ran_on_gpu_says_nothing` stays green — so the tests discriminate rather than just failing. `used_gpu: None` is Unknown, not a fallback: refusing on it would make every non-reporting path a hard error. NOT included, deliberately: the rejection's REASON (cosine, position) is on stderr but not in the JSON. It is produced inside realizar's F2 gate and no channel carries it to the CLI; adding one is a #3606-shaped follow-up, not something to approximate with a guess here. Refs #3602, #3483 Pmat-Ticket: PMAT-3602 * test(PMAT-3346): the measured Qwen3.5-0.8B inventory the dense path cannot reproduce QE2E-INV-001 could not be judged because nothing in the tree held a MEASURED Qwen3.5 tensor inventory to judge against. This adds one: the 320 tensors of ~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782...dc517), read straight from the GGUF header rather than from a model card. It already falsifies the current arithmetic. Dense/GQA accounting applied to that file gives 644,400,128 against a measured 752,393,024 — short by 107,992,896, 14.4% of the model, because 18 of the 24 layers are Gated DeltaNet and no term here counts their conv, gate, state or output projections. Two shapes in the file are also not what dense accounting predicts, and both are pinned: attn_q is [1024, 4096] = 2 * num_heads * head_dim (the q projection emits the attention output gate alongside the query; attn_output [2048, 1024] confirms num_heads * head_dim = 2048), and attn_q_norm/attn_k_norm are present at head_dim. The file is also TIED — it has no output.weight. Refs #3346 Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(PMAT-3346): ModelConstraints carries the gated-DeltaNet shape, so a hybrid model can be counted contracts/model-families/qwen3_5.yaml declares inner_size, state_size, conv_kernel, group_count and full_attention_interval under constraints:, and ModelConstraints carried none of them. A Gated DeltaNet layer's parameters live entirely in those dimensions, so every consumer of the descriptor counted Qwen3.5 as if three quarters of its layers did not exist. Carried through as ModelConstraints::deltanet: Option<DeltaNetShape> — the runtime YAML loader (parsing.rs) and the compiled-in registry (build_parsing.rs + build_codegen.rs) both populate it, and FALSIFY-MF-QWEN35-010 pins that the declared values survive the trip and that no other family acquires a shape it never declared. qwen3_5.yaml is the only descriptor with these keys, so every other family keeps byte-identical accounting. model_arithmetic gains gated_deltanet_layer_params (one term per GGUF tensor: attn_qkv, attn_gate, ssm_conv1d, ssm_alpha/beta, ssm_a, ssm_dt.bias, ssm_norm, ssm_out) and hybrid_layers (the interleaved schedule). attention_layer_params gained two terms the real file has and dense accounting did not model: the gated q projection (2*n_h*d_k) and the q/k norm vectors. Falsified against a real model, not against itself: fed the 0.8B configuration, the equation now reproduces the 320-tensor inventory of Qwen3.5-0.8B-Q4_K_M.gguf EXACTLY — 752,393,024, both layer kinds matching tensor for tensor. QE2E-INV-001 is still NOT asserted, and no range was widened to make it pass. The 9b descriptor now yields 8,344,907,136, up from 8,208,519,168 but still 0.655B below [9.0B, 9.2B]. The remaining gap looks like descriptor drift rather than missing arithmetic: 9b keeps inner_size 2048 — the value the 0.8B uses at hidden_dim 1024 — while quadrupling hidden_dim, and its group_count 8 fails 8 * 128 == 2048, a consistency the measured 0.8B satisfies at 16 * 128. Only a real Qwen3.5-9B file can settle it; none is on this box. Refs #3346, #3347 Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * contracts(binding): the QE2E-INV-001 note blamed a gap that is now closed The note said the obligation was undischarged because ModelConstraints does not carry the DeltaNet shape keys. It does now, and the 0.8B count reproduces the real GGUF exactly. What actually blocks the obligation is descriptor drift at the 9b variant. Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3346 (adoption): the QE2E-INV-001 note lost its closing quote, and pv extract shrank the graph by 356 triples instead of refusing d3cc76f1f rewrote the note on the QE2E-INV-001 binding and dropped the trailing `"` — contracts/binding.yaml stopped being valid YAML at line 764 (`found unexpected end of stream`). Nothing in the PR noticed because `pv extract contracts` does not refuse a binding registry that will not parse: it emitted a graph with 15,244 triples where main has 15,600 — every bound symbol AFTER the broken entry (prune::run, distill::run, harness_ir::*, ptx_explain::run, …) silently gone — and `--check` would have agreed with itself. Found while regenerating the derivative for this adoption, by the drop, not by any gate. One character. With it, binding.yaml parses (156 entries, same as main) and the extraction is byte-identical to main's committed contracts.nt, so this PR owes no graph change after all. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(roadmap): mint PMAT-3602 as a work item so the AD-04 quorum can resolve it `quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work` reads the aggregated roadmap rather than an id counter. The branch carried the trailer `Pmat-Ticket: PMAT-3602` while no such work item existed — a GitHub issue number is not a work item. The id is DERIVED from the issue (`pmat work add --github-issue` calls that "the path to prefer": GitHub allocates centrally, so two agents cannot be handed the same id, which `max(id)+1` cannot promise). Fragment, not a direct edit to roadmap.yaml: additive, one file, and it does not make every stacked branch dirty on the same lines. Acceptance is hand-entered from the issue's done_when — `pmat work add` derives neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414). Guards: diff_additive, fragment_required, ids_unique, sorted, completion_is_cited all exit 0; `make roadmap-aggregate-check` reports `roadmap.yaml == aggregate(48 fragment(s)), idempotent`. Refs #3602 Pmat-Ticket: PMAT-3602 * roadmap: PMAT-3637 — the registration row this PR is, so its quorum judges fidelity not implementation Round 0 ran with --ticket PMAT-3604 and all three lanes FAILed for the same reason: the diff implements none of PMAT-3604's criteria. Correct — it was never meant to; it registers the ticket so #3634's quorum can start. The receipt is kept as quorum-PMAT-3637-r0-misframed-as-3604.json (a FAIL is a record, not a mistake to erase). PMAT-3637 is the row this diff satisfies: fidelity of the PMAT-3604 entry to issue #3604's done_when, additive aggregate, nothing outside docs/roadmaps/. Round 1 runs against it. Refs #3604, #3634 ont-delta: none — roadmap entries only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(roadmap): mint PMAT-3605 as a work item so the AD-04 quorum can resolve it `quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work` reads the aggregated roadmap rather than an id counter. This branch carried the trailer while no such work item existed — a GitHub issue number is not a work item until something mints it. The id is DERIVED from the issue: `pmat work add --github-issue` calls that "the path to prefer", because GitHub allocates centrally and `max(id)+1` cannot promise two agents different ids. A fragment rather than a direct roadmap.yaml edit — additive, one file, and it does not make every stacked branch dirty on the same lines. Acceptance hand-entered from the issue's done_when; `pmat work add` derives neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414). Guards: diff_additive, fragment_required, ids_unique, sorted, completion_is_cited all exit 0; `make roadmap-aggregate-check` idempotent. Refs #3605 Pmat-Ticket: PMAT-3605 * PMAT-3346: the roadmap row, so the AD-04 quorum for #3350 has a ticket to judge against Acceptance transcribed from issue #3346 as this PR answers it (the type that reads the descriptor was wrong, not the range or the descriptor), with the 9B range instantiation explicitly out of scope until a real 9B GGUF exists. Refs #3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * roadmap: round 1 said two true things — fix both Lanes 1+2 (gemini-3.1-pro-high, gemini-3.8-flash-high) FAILed PMAT-3637 on (a) the round-0 receipt committed under docs/audits/, which criterion (3) forbade — dropped; the receipt judged the wrong question and its ticket field would read as a verdict on PMAT-3604; and (b) a criterion in the PMAT-3604 notes that issue #3604 does not state ('Only an Accepted verdict writes a receipt') — it is PR #3634's design decision, not a done_when item; removed from the transcription. Criterion (3) now admits this row's own receipt at docs/audits/quorum-PMAT-3637*.json, which it must, or the PASS receipt could never be committed. Refs #3604, #3634 ont-delta: none — roadmap entries only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(run): quorum round 1 found the --json surface skipped and a branch that could never run Two findings from lane 1 (gemini-3.1-pro-high) of the AD-04 quorum, both correct, both fixed. Lane 3 passed the PR without seeing either. 1. `reconcile_accelerator(...)?` ran BEFORE `print_run_output`, so a rejected `--gpu` run early-returned and `--json` emitted nothing at all — the exact surface #3602 item 1 names. Machine surfaces (`--json`, `--stream`) now emit before the refusal propagates; the human surface still prints nothing. This is a DELIBERATE deviation from `after_generation`'s contract, which says the caller "must print NO output" on a forced refusal. Named rather than quiet: that rule exists so a CPU result is never read as a GPU success, and a document carrying `"backend": {"fell_back": true}` beside exit 14 cannot be read that way, while a human-formatted success blob can. The protective half is kept; the half that blinded `--json` consumers is not. 2. The caller's `if let Some(note)` arm was UNREACHABLE. `announced` was `Some("gpu")` exactly when `forced` was true, and `after_generation`'s corrective-line branch requires `forced == false` — so the Some arm could never execute. A branch with no reachable caller, which is precisely the defect this PR fixes, reproduced one layer down while fixing it. `reconcile_accelerator` now returns `Result<()>`, wires the FORCED half only, and says in its doc comment why the default-selection half is not wired: nothing calls `registry::announce`, so there is no recorded announcement for a default selection, and manufacturing one would assert a choice this process never made. That is REG-8, not this PR. Verification: cargo test -p apr-cli --lib = 7289 passed, 0 failed, 12 ignored; cargo fmt --all --check = 0; cargo clippy -p apr-cli --lib -D warnings = 0. Refs #3602 Pmat-Ticket: PMAT-3602 * PMAT-3346 (adoption): a bias vector is as wide as the projection it biases Found by the AD-04 quorum on #3350 (lane 1, gemini-3.1-pro-high, cited model_arithmetic.rs:144): the projection term used q_out for a gated family's q matrix (2*n_h*d_k — attn_q emits the output gate, MEASURED in Qwen3.5-0.8B) while the bias term still used q_dim. A bias narrower than its projection is not a model. One token: q_dim -> q_out in the bias sum. Why a delta-0 measurement did not catch it: no shipped family exercises the case. Qwen3.5 has no attention bias; Qwen2.5 has biases but is not gated, so q_out == q_dim there. The test that pinned 12 (a q_dim bias under a q_out matrix) now asserts 16 and says why, and a second test holds the other polarity — a non-gated family with biases is unchanged at 12. 115 model_arithmetic + model_family tests pass; oracle 218; clippy clean. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(run): quorum round 2 — a --json --benchmark leak, a comment claiming coverage it lacks, and a stray artifact I swept in Three findings, two lanes, all correct. 1. LANE 1 — `--json --benchmark` leaked a human success blob on a refused run. The refusal path carried its OWN copy of "is this a machine surface" (`stream || output_format == "json"`) and dropped `!benchmark`. So that combination entered the branch, matched neither machine arm inside `print_run_output`, and fell through to the human benchmark rendering — for a run being refused. Two spellings of one condition drifting apart in the gap between them, which is the defect this PR's own dispatch comment warns about. One spelling now: `emits_machine_output(stream, output_format, benchmark)`. `the_machine_output_predicate_matches_print_run_output` pins it to the arms it describes over every flag combination; planting the pre-fix spelling turns it and `json_plus_benchmark_is_not_a_machine_surface` RED, verified. 2. LANE 1 — the classifier comment claimed `--gpu-layers all|n` was among the flags handled here. It is not: `Commands::Run` carries `gpu` and `no_gpu` and nothing else; `--gpu-layers` belongs to `apr serve`. So `layers_want_accelerator: false` is CORRECT and the comment was the defect — a comment asserting coverage the code does not have. Both comments now say the input is genuinely absent for this surface rather than stubbed. 3. LANE 3 — `trace-1789944278.json` (232 lines) was committed to the repo root. A runtime byproduct of `test_print_chrome_trace_creates_file`, which writes `trace-<epoch>.json` into the CWD when given no `--trace-output`. I swept it in with `git add -A` immediately after a quorum round — the exact thing my own notes say never to do there. Removed, and this commit stages files by name. The test writing into the repo root is a separate defect and is not fixed here. Round 2 was 2 FAIL / 1 NO-VERDICT. Lane 2 has now returned NO-VERDICT twice with `envelope_status: SUCCESS`, `transport_status: SUCCESS` and non-empty `raw_bytes` — it answers, and the harness cannot extract a verdict from what it returns, so width 3 has been delivering two votes. Verification: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored; cargo fmt --all --check = 0. Refs #3602 Pmat-Ticket: PMAT-3602 * roadmap: transcribe #3604 verbatim — round 2's pro lane was right Criterion (5) had grown an implementation-status clause from PR #3634's report ('the --json field lands with #3606 StageTimings…') that issue #3604 does not state, and the out-of-scope list was a paraphrase. PMAT-3637 (1) says every criterion present, none added: the notes now carry the issue's done_when 1-6, admission rule and out-of-scope list as written. Refs #3604, #3634 ont-delta: none — roadmap entries only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * audit: quorum-PMAT-3637 — 3/3 PASS, round 3 on the ruled trio pro / 3.7-flash / 3.6-flash, three conversations, no silences. Rounds 0-2 each said something true about this PR: round 0 judged the wrong ticket; round 1 found a receipt outside scope and an invented criterion; round 2 found an implementation-status clause in a 'none added' transcription. None is committed — each judged a diff that no longer exists. Refs #3604, #3634 ont-delta: none — audit artifact only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3346: quorum verdict 3/3 on e060c7e2c (AD-04) Round 0 was 1 FAIL / 1 no-verdict / 1 PASS and the FAIL was real (the bias width, fixed in e060c7e2c). Round 1 on the fixed head: 3/3 PASS, gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(audits): AD-04 quorum receipt for PMAT-3577 — AGREED 3/3 Three independent agy lanes on distinct models, none in the author's family (author measured as Opus 5 / claude): gemini-3.1-pro-high = PASS gemini-3.7-flash-high = PASS gemini-3.6-flash-high = PASS Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG rather than by editing the shared config, which three sessions were launching against concurrently. THE MERGE DID NOT CHANGE WHAT WAS JUDGED. The lanes ran against 7f82e2cd3; the PR head is 4fbd2ec31, a merge of main. The trees differ by four files, but all four arrived FROM main, and the PR's own contribution relative to main is byte-identical across the merge: files, pre-merge : 49 files, post-merge : 49 only in post : (none) only in pre : (none) sha256(diff vs merge-base), pre-merge : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35 sha256(diff vs merge-base), post-merge : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35 So this is the same head for review purposes and does not need a new round. A naive `git diff origin/main..HEAD | sha256sum` from my local head disagrees only because that head carries this receipt commit, which the lanes never saw — comparing the receipt's parent is what makes the two comparable. Refs #3600 Pmat-Ticket: PMAT-3577 * fix(tests): the shape-count ratchet was red — #3605 adds a 6th shape (quorum round 1) `the_tracked_repo_graph_is_fresh` asserts `shapes_n == 5` and names the five. `refusal-receipt-v1` makes six, so this PR shipped a CI red that a quorum lane found and I did not. MEASURED, not guessed: `pv extract contracts --check` on this branch reports `shapes_n: 6, triples: 15620`. The hardcoded count stays hardcoded. A shape added without anyone noticing is exactly what this assertion exists to prevent, so adding one is SUPPOSED to turn it red and make you name the new shape. Noted in the comment that the counter is shared across branches — a sibling PR adding a shape (#3600's parity-receipt-v2) will need it raised again at merge, which is the ratchet working rather than a conflict to route around. Refs #3605 Pmat-Ticket: PMAT-3605 * chore(audits): AD-04 quorum receipt for PMAT-3602 — AGREED 3/3 gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS. Author measured as Opus 5 (claude); no lane in the author's family. Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared config that other sessions were launching against. This trio was chosen on measured odds after flash-class lanes returned NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my earlier runs). All 15 lanes voted in this batch. Refs #3638 Pmat-Ticket: PMAT-3602 * chore(audits): AD-04 quorum receipt for PMAT-3605 — AGREED 3/3 gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS. Author measured as Opus 5 (claude); no lane in the author's family. Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared config that other sessions were launching against. This trio was chosen on measured odds after flash-class lanes returned NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my earlier runs). All 15 lanes voted in this batch. Refs #3613 Pmat-Ticket: PMAT-3605 * PMAT-3351 (adoption): re-id the fragment — PMAT-3347 is #3348's row on main Both #3348 (merged; bound the qwen35-e2e equations, closed #3347) and this PR (the L2 column reads a link, not an index) address issue #3347, and both claimed roadmap id PMAT-3347. Two PRs cannot share a row: main's PMAT-3347 stays as #3348's, this PR's fragment becomes PMAT-3351 (github_issue 3347, same title, same notes), and the aggregate is rebuilt from main's roadmap.yaml plus this branch's fragments — 935 + 1 = 936, nothing re-serialised. Refs #3347, #3348 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3351: notes transcribe issue #3347's Ask verbatim (bullets 2 and 3; bullet 1 was #3348) The quorum lanes read pmat work status, i.e. this fragment's notes, not the GitHub issue. Empty notes would have them judge against a title. The issue has no done_when section, so its Ask and its two defect statements are transcribed as written, with the out-of-scope bullet named and no count offered as a criterion. Refs #3347 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3338: register the second ticket this PR closes, so the quorum judges the truncate fix as in scope Quorum round 0 on a4b76397f was 2 FAIL / 1 PASS, both FAILs on one point: the branch's first commit fixes #3338 (--table panicked on a byte-index cut inside a multi-byte char) and PMAT-3351 does not ask for it. Both lanes called the in-scope work correct. The fix is a prerequisite — the L2 column is exercised through --table on the real corpus, which panicked before it — and #3338 is an OPEN issue this PR genuinely closes. So it gets its row, transcribed from the issue, and round 1 runs with --ticket PMAT-3351,PMAT-3338. Refs #3338, #3347 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3604: the receipt's temp file is private to its writer — a shared name made the atomicity claim false AD-04 quorum round 0 on #3634, lane 1 (gemini-3.1-pro-high), cited f2_receipt.rs:180: write_receipt used one shared <sha>.json.tmp, so two apr runs validating the same model at once could have writer B truncate the file writer A was about to rename, and A rename B's partial into place. The reader would classify it Unreadable and validate — the safe direction, done_when 4 — but the docstring said 'never a truncated one', and that was false. The temp name now carries the pid and a per-process counter; a failed rename removes its own temp. A racing test spawns two writers on one path forty times and parses the survivor each time: it is always one writer's WHOLE receipt. 14 tests. Lane 1's other two findings, for the record: --revalidate is done_when 2 verbatim (the fragment carries it), not scope creep; and F2Outcome's fields are the values the stderr line already prints and done_when 5's data path for #3606's JSON, not dead code. Two lanes passed the same diff. Refs #3604 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3604: quorum verdict 3/3 on dd546592b (AD-04) Round 0 was 2 PASS / 1 FAIL; the FAIL was the shared temp-file race, fixed in dd546592b. Round 1 on the fixed head: 3/3 PASS, gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent. The PMAT-3604 roadmap row was on this checkout UNCOMMITTED for the resolver (pmat work status reads the checkout); it lands via #3637. The lanes judged the committed diff origin/main...HEAD. Refs #3604, #3637 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3351 + PMAT-3338: quorum verdict 3/3 on d915b8f16 (AD-04) Round 0 (--ticket PMAT-3351 alone) was 2 FAIL / 1 PASS, both FAILs on the #3338 truncate fix being out of scope; both called the in-scope work correct. Round 1 with both tickets registered and named: 3/3 PASS, gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent. Refs #3347, #3338 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(contracts): the v1 contract had no equations, and its own denominator guard was never wired Two reds on this PR, both mine, and the second is the better finding. 1. `contract_data_integrity` is a SHRINK-ONLY ratchet at 444 and this branch made it 445. The 445th is `parity-receipt-v1: no equations`. Measured which one rather than guessed — `parity-receipt-v2` is not in the list. `kind: pattern` correctly declares no SHAPE (a shape over a class nothing instantiates passes vacuously, which is #3610's whole lesson), but equations are not shapes, and this contract does assert things: PRC1-INV-001 and PRC1-INV-002 already said them. They are now written as the two equations the integrity check reads — `no_instances` and `refused_by_name` — rather than as new claims invented to satisfy a counter. 2. `check_guards_are_wired.sh`: `NEW: parity_receipt_denominator.sh`. I ADDED A GUARD IN THIS PR AND WIRED IT INTO NOTHING — the second, independent reader of the receipt corpus, shipped where no workflow names it. Minutes before finding this I wrote into that same contract's equations, as a precondition: "both readers are wired into CI — a refusal nothing runs is not a refusal." My own precondition, violated by my own PR, in the same file. `check_guards_are_wired.sh` caught what I had just finished writing down. Now named in `ci.yml` beside the other cargo-free, model-free case tables. Its 4-case self-test passes: a legacy record refused by name, a receipt added without bumping the denominator disagreeing, bumping it making them agree. NOTE — THIS TOUCHES `.github/workflows/ci.yml`, which CLAUDE.md lists as a check-in item. Additive only: one step in an existing guard block, no matrix, trigger or gate logic changed. Wiring an unwired guard is the minimum the failing check asks for, and leaving it unwired to avoid the file would be keeping a refusal that never runs. VERIFICATION cargo test -p aprender-contracts --test validate_contracts contract_data_integrity 1 passed bash scripts/parity_receipt_denominator.sh --self-test rc 0, 4/4 bash scripts/check_guards_are_wired.sh PASS (4 -> 3) pv validate contracts/parity-receipt-v1.yaml 0 errors, valid DIFF MOVEMENT, for the receipt: the judged diff DID change — an equations block and one CI step. No behaviour the lanes reviewed was altered; the extractor, the shapes and the denominator script are byte-identical. AD-04 re-review is the inspector's call. Refs #3577 Pmat-Ticket: PMAT-3577 * PMAT-3401 (adoption): the three derivatives 24 new contracts oblige — census, graph, README — regenerated together This PR adds 24 contracts and regenerated none of the tracked artifacts derived from the corpus. On #3581 that omission surfaced one per CI round, each masked by the one before it. All three here, at once, with a pv built from this tree under a pinned target dir: contracts/census.json 1800 -> 1824 (+24, the contracts added) contracts/contracts.nt 15,600 -> 15,696 triples (+96 = 24 x 4; GREW — a drop is the tell for a malformed input or a stale binary; binding.yaml still parses, 156 entries) README CONTRACT_COUNT 2 blocks -> 1824 via make readme-sync The merge took main's generated README blocks over the branch's hand-typed 1866, then readme-sync wrote the measured 1824; the branch's number described a tree that never existed on main. Verified: all 24 contracts pv-validate under the pinned binary (control: main's crux-A-01 valid under the same one); lint_passes_on_real_contracts green, so the sigma prose ratchet holds; 1666 engine tests; ont4b shapes gate 11/11; test-binding ratchet; readme-sync-check; FALSIFY-README-002; roadmap additive added=26 deleted=0, aggregate idempotent, 26 fragments present. Refs #3401 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(run): extract reconcile_and_emit — run() hit cognitive 27 against a 25 ceiling `check_complexity_ratchet.sh` went RED with: RED NEW crates/apr-cli/src/commands/run_entry.rs::run cyclomatic 21 cognitive 27 (over a threshold, absent from the comparand) The reconcile-then-emit block I added in this PR is what took `run` over. Moved into `reconcile_and_emit`, which also gives the ordering decision — machine surfaces emit on a refusal, the human surface does not — a place to be documented that is not the middle of a 30-argument entry point. Now: PASS (D2): e6f77c98c vs 19750b384 measured by pmat 3.41.1 — none new, none grown. NOTE FOR THE NEXT PERSON: I first re-ran the ratchet against my WORKING TREE and saw the identical cognitive 27, and briefly concluded the extraction had not helped. It had. The script measures two REVISIONS — it prints `merge HEAD <sha>` — so an uncommitted fix is invisible to it. Commit, then measure. Tests: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored; fmt --check = 0; clippy -D warnings clean. Refs #3602 Pmat-Ticket: PMAT-3602 * PMAT-3496: a registration row for the registration PR — the author's earlier receipt named this ticket but no row existed on the branch * PMAT-3496: clause (1) transcribes issue #3495's title verbatim (132 Kani harnesses; Verus pilot), not a paraphrase * fix(guard): a step NAME read as a bare guard_tree.sh dispatch hid four dark guards; the #3305 guard runs nowhere and dies on intel's python 3.10 (#3644) Refs #3644 #3646 #3305 #3626 THE FINDING, proven by mutation before it was argued. check_guards_are_wired.sh counts a guard wired-by-dispatch when a workflow runs guard_tree.sh in a mode whose --dry-run RUN set contains it. Its invocation regex accepts `guard_tree.sh` followed by `'` (for a quoted run: string). ci.yml has two step NAMES, `"guard_tree.sh's own case table …"` and `"guard_tree.sh's parallel dispatcher …"`; the apostrophe matched, neither line says --no-cargo, so each was read as a bare dispatch in mode "all" -- 55 cargo-classified guards counted wired that `guard-tree` (--no-cargo) never runs and `guard-cargo` never names. Rewriting those two names (nothing else) turns the meta-guard RED: 3 -> 7 unwired -- check_model_ladder.sh, check_pathonly_devdeps_unused_in_src.sh, check_pr_review_counts.sh, check_pr_review_receipt.sh. Three of the four are cargo-classified by a COMMENT that says the build tool's name; none of those three invokes it. check_pathonly_devdeps_unused_in_src.sh (#3305: src/ may not use a dev-dep that publishing deletes; clean-room red 8/8 across two releases) has therefore run nowhere since it was written on 2026-09-15. And it could not have run on the intel hosts: `import os, re, sys, tomllib` at module level, python 3.10.12 there with neither tomllib nor tomli (measured on mac-server tonight; the sovereign-ci container has no python3 at all). Reproduced with python3.10 -S: the traceback is swallowed by `out=$(scan … || true)` and the case table prints "FAIL: missed hit_use" ×4 -- the regex reported broken by an interpreter that never ran it. Same class as #3626's guard on intel-clean-room-6. WHAT CHANGES scripts/check_guards_are_wired.sh - `_not_a_name_line` drops `name:` lines before the invocation match, in dispatcher_wired() and the per-guard scan. A step name is documentation whatever it contains. - rows 8-10: the exact ci.yml fixture (`- name: "guard_tree.sh's own case table …"`, no run: line) must leave the dispatched guard unwired; a QUOTED run: line still dispatches (the control the `'` exists for); a step named after a guard wires nothing. Mutant (drop removed): rows 8 and 10 RED. - the ledger's ratchet is now set-aperture, owned by this file. scripts/lib_baseline_ratchet.sh, scripts/check_baseline_ratchets.sh - set-aperture gains a NAME-entry admission for sets of FILES: an added entry with no `:` is admitted iff the comparand carries that file (at the entry's path, or beside the owning guard -- the ledger names siblings by basename and guard_tree.sh reads it so) AND the owning guard changed in the diff. A file this branch created is refused (PERF-028's shape). Six rows: predates -> admitted; branch WROTE it -> refused; no guard edit -> refused; escaping path -> refused; beside the OWNER -> admitted; beside a different owner -> refused. Lib mutants (always admit / branch removed / sibling resolution removed) each turn rows RED; a separate absolute-path check survived its mutant because `git cat-file -e` already refuses those, so it is not there. - classify: unwired_guards_baseline.txt -> set-aperture, owner named. scripts/unwired_guards_baseline.txt 3 -> 6, as APERTURE REVEALS with a reason per line (each may only leave): check_model_ladder.sh (T-2 release gate, the multiplatform_dogfood class); check_pr_review_counts.sh (RED on main today, "6 row(s) disagree" -- #3646); check_pr_review_receipt.sh (takes a receipt path; nothing passes one -- #3646). check_pathonly_devdeps_unused_in_src.sh is NOT ledgered: see below. scripts/check_pathonly_devdeps_unused_in_src.sh - readers tomllib -> tomli -> ENV rc=2 naming python version and $RUNNER_NAME. No purpose-built manifest reader: a second TOML implementation over 78 manifests, where a wrong "has source" is a silent PASS on exactly the #3305 class, is not a fallback worth shipping. UNMEASURED-with-a-name is the honest state on a runner with no TOML library. - the scanner ends with `SCAN-DONE manifests=N reader=…`; the shell REQUIRES it. Output without it (no reader, /bin/false, no python3) is ENV rc=2 -- never rc=1, never "no violations". The selftest and the tree scan propagate 2 instead of swallowing it. - five death rows: both readers absent (sys.modules nulled so the imports fail for real); interpreter exits 1 silently; no interpreter (127); the working interpreter as control; and THE WHOLE GUARD under the intel shape (nested, recursion-guarded) -- the call site is where `|| true` hid it. Mutants: trailer not required (3 rows RED); trailer printed on the no-reader path (1 RED); `|| true` restored (the whole-guard row RED). - three comments and one FAIL string reworded so the file no longer says the build tool's name with a space after it: it invokes none, and that substring is what CARGO_RE classifies on. guard_tree.sh --dry-run --no-cargo now lists it as `run:`; bare run on this tree rc=0 in 3 s. Passes: check_guards_are_wired --self-test 10/10; check_baseline_ratchets --self-test 61 rows; pathonly --selftest under python 3.13.1 and 3.10.12 (tom…
…ate turns down, instead of "parity disproven" (5) (paiml#3689) * fix(pv): --table panicked on the real corpus — it cut a property by BYTE index (#3338) `pv proof-status contracts/ --table` exits 101 on `contracts/`: byte index 40 is not a char boundary; it is inside '∈' (bytes 39..42) of `for Q4_K_M Qwen2.5-Coder, quantization ∈ {Q4_K, Q6_K}` The column width is a byte count (`property.len()`, capped at 40) and `truncate` sliced `&s[..max]`, so any property whose byte 40 lands inside a multi-byte char panics. Eight contracts in `contracts/` do; the first one walked is `apr-inspect-quantization-v1.yaml`. The budget stays a byte budget — the table is laid out in bytes — and the cut now walks back to the nearest char boundary. Two tests, both RED before this commit (each panicked at obligation_matrix.rs:169): the helper row, and one through `format_obligation_table` because that is the path the operator hit. The helper fixture asserts `!s.is_char_boundary(40)` first, so it cannot silently stop proving anything. Pmat-Ticket: PMAT-3347 Refs #3347, #3338 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pv): the L2 column read an INDEX, not a link — so 3,573 of 3,753 obligations ticked (#3347) `obligation_matrix` computed `l2_tested` as `idx < falsification_tests.len()`. Obligation 3 was "tested" because the contract happened to hold 4 tests, whoever those tests were about. All 7 obligations of `qwen35-e2e-verification-v1` showed ✓ before a single test existed. The fallback was a substring match between the obligation's `property` prose and a test's `rule` prose, which is an inference, not a claim. WHICH LINK, decided by counting the corpus (1,842 files, 3,792 obligations, 4,691 falsification tests), not by preference: proof_obligations[].discharged_by 89 (62 `falsification_tests[N]`, 13 a test id, 1 a YAML sequence, rest comma lists / prose / a kani id) falsification_tests[].binds_to 38 falsification_tests[].obligation 26 (12 name an obligation id, 6 the exact property text, 8 dangle) proof_obligations[].id 180 of 3,792 kani_harnesses[].obligation 1,918 — that is the L3 column, not this one All three L2 spellings are DECLARATIONS by the contract author, so all three are read. `binds_to` is a serde alias of `obligation`, which is safe only because no entry carries both keys — checked, because serde turns that into a `duplicate field` parse error rather than a silent pick. Not read: `applies_to`. 12 `binds_to` values match one, but `AppliesTo` is an enum with `#[serde(other)] Other`, so the string is discarded at parse and there is nothing left to compare. Those 12 report `?`, not a false ✗. WHAT THE COLUMN NOW SAYS. Three values, because "no test covers this" and "nothing here says which test covers what" are different facts: ✓ Tested a test in this contract cites this obligation ✗ Untested the contract's links resolve and none names it — or it ships no falsification test at all, which is a reading, not a gap ? Unknown no readable link; not measured. An unread window is Unknown, never a tick MEASURED over `contracts/`, and the drop IS the point — it is what the old column was hiding: before after L2 ✓ 3,573 86 L2 ✗ 180 65 L2 ? 0 3,602 (3,753 obligation rows, 873 contracts) Nothing was adjusted to keep the number up, and no threshold was added. The two link fields did not exist on the structs, so both keys were written to disk and silently dropped on parse — the shape of #3314 (`id`) and #2465 (`test_harness`). `discharged_by` is typed `Citation` (scalar | comma list | YAML sequence) because `Option<String>` failed the WHOLE corpus on `publish-manifest-v1`: `invalid type: sequence, expected a string`. RED first: `l2_does_not_tick_for_an_obligation_no_test_cites` — two obligations, two tests, both citing OB-A. It asserted through the rendered table (the surface that was lying, and API-stable), so it is the SAME test before and after: on the old code `Beta holds | ✓`. Its OB-A arm keeps a fix that merely stopped ticking everything from passing. Out of scope_paths, and forced: `lint/strict_test_binding.rs` holds the only EXHAUSTIVE `FalsificationTest` literal in the tree (no `..Default::default()`), so no schema field can compile without that one line. Pmat-Ticket: PMAT-3347 Refs #3347, #3091, #3114 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(pv): single-file --strict-test-binding reported every ref missing — it now refuses (#3347) The gate resolves cited test names against a source index rooted at the contract path's PARENT. The directory form gets the repo root and finds `crates/`; the single-file form gets `contracts/`, which holds no source, so every cited ref resolves to nothing and all of them are reported missing. Measured on `contracts/pv-artifact-kinds-v1.yaml`, a control whose eight refs all resolve: pv lint contracts/pv-artifact-kinds-v1.yaml --strict-test-binding total_refs 8, existing 0, missing 8 pv lint contracts/ --strict-test-binding total_refs 548, existing 521, missing 27 <- none of the 27 is this one A control contract failing identically to a broken one is a gate that cannot discriminate, and it fails SILENTLY: `passed` is true in non-strict mode, so the run still says `Result: PASS` while printing eight false findings. REFUSED rather than repaired. The scan root is computed in `provable_contracts::lint::run_lint`, outside this ticket's scope; a refusal lives in the caller, is honest, and cannot be mistaken for a clean bill. If the root is later made explicit (a `LintConfig` field the CLI can feed, which is what `--crate-dir` does NOT do today), this refusal is what should be deleted. Exit 1, not the exit-2 `decline:` class: exit 2 belongs to `ZeroContracts`, whose message ("0 contracts under ...") would be false here — there IS a contract; it is the gate that cannot run over it. Two tests, in the CI-wired `cli_integration` target: the refusal names the flag and prints no PASS and no findings, and a directory-form control proves the refusal is specific to the single-FILE form rather than to the flag. Pmat-Ticket: PMAT-3347 Refs #3347 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(crux): AutoGluon becomes a CRUX competitor — category O, 24 contracts, 25 tickets, registry edit + mutation proof Research of ../autogluon (1.6.3 @ 77946149) as a competitive-research source for aprender, landed the way category N landed linfa and burn (#3169): the competitor is admitted to the CLOSED registry, every story is a real contract with falsification gates, every story is a registry row, and every missing/partial story has a GitHub issue and a roadmap fragment. What AutoGluon is, measured from the tree rather than recalled: three predictors. TabularPredictor (65 public methods, 11 presets, 24 model families of which 8 are tabular foundation models added in 1.4-1.6), TimeSeriesPredictor (Chronos-2/Toto-2 pretrained, 30+ local/deep models, 16 metrics incl. WQL/MASE/RMSSE, auto backtesting since 1.5) and MultiModalPredictor. Evidence under evidence/crux/autogluon/. What aprender has, measured at eb262f8eb: automl/ is a single-estimator hyperparameter tuner (TPE, grid, random, DE, TimeBudget, EarlyStopping); time_series/ is one univariate ARIMA; encoders, calibration, SHAP/LIME/ permutation importance and KFold/cross_validate exist as building blocks. No predictor-level fit(label), no leaderboard, no bagging, stacking or greedy weighted-ensemble selection, no panel forecasting, no quantile forecast metrics. The gap is the AutoML UX, not the algorithms. Category O — AutoML Parity — 24 stories: 9 P0 (the README hello-world: fit(label), problem-type inference, presets, leaderboard, feature pipeline, weighted ensemble, time budget, panel forecaster, quantile metrics), 9 P1 (bagging, stacking, importance, threshold calibration, deployment artifact, tabular foundation model, backtesting, local baselines, pretrained forecaster), 6 P2 (refit_full, distill, infer_limit, fit diagnostics, memory-aware fit, covariates). MultiModalPredictor, autogluon.cloud, MLZero and Ray-parallel fits are CUT on the epic with reasons. CRUX_COMPETITORS: [&str; 14] -> [&str; 15] + autogluon Not a BEAT pillar: aprender claims no pinned-benchmark win over AutoGluon. Both registry tests that keep BEAT_INCUMBENTS and CRUX_COMPETITORS apart are extended, not worked around. Tickets: epic #3370, stories #3371-#3394, label pareto-autogluon. Roadmap: 25 fragments under docs/roadmaps/entries/, roadmap.yaml regenerated by the aggregator (idempotent check passes). Spec: docs/specifications/crux-competitive-research-ux-workflows.md v2.2 -> v2.3 — §3 gains rows for linfa+burn (category N, which #3169 never recorded there) and AutoGluon; §5 gains Category O; §6 notes that coverage_intake in the YAML is the source of truth. coverage_intake 267 -> 291 (partial 72 -> 77, missing 152 -> 171). Verification: - pv built from THIS tree validates 25/25 (24 new + master). The stale ~/.cargo/bin/pv rejects crux-O-01 with CRUX-002 — the behavioural delta proves the registry edit engaged. - Mutation-verified: deleting "autogluon" from CRUX_COMPETITORS turns competitor_registry_covers_the_corpus_vocabulary RED with "autogluon is used by contracts/ and must stay in CRUX_COMPETITORS" and the_real_crux_registry_rows_are_all_in_domain RED. Restored: 20/20. - cargo test -p aprender-contracts --lib: 1526 passed, 0 failed. - Every falsification gate is LIVE-PENDING prose (no `::`), so strict-test-binding has nothing to refuse; the obligations are RECORDED as unfalsifiable-by-absence, not satisfied. - README CONTRACT_COUNT regenerated 1835 -> 1866 by readme_sync.sh. - Guards: roadmap fragment/ids/sorted/additive/completion, contract test-binding and enforcement, shell-lint ratchet, hardcoded paths, readme claims, grep -q ratchet — all rc=0. Pmat-Ticket: PMAT-3370 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(crux): bind the admission PR to its own ticket PMAT-3401 — the epic's acceptance criteria describe the programme, not this diff Round 2 of the quorum read the epic (PMAT-3370) as the ticket and refused the admission for not implementing the 24 stories. The admission is its own unit of work with its own done-when; this fragment says so. Closes #3401 Pmat-Ticket: PMAT-3401 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(crux): PMAT-3401 title inventories the diff exactly — 26 fragments (its own included) and the scaffold CATEGORY_NAMES edit Quorum round 3 lane 1 refused on two literal mismatches between the ticket and the diff: '25 roadmap fragments' (there are 26 once this ticket's own fragment lands) and an unlisted edit to scripts/crux_scaffold_contracts.py. The title now lists every path. Pmat-Ticket: PMAT-3401 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * roadmap(PMAT-3495): VERIFY-001 on 0.71.0 — run the 132 Kani harnesses in CI (proof credit from runs, not declarations), then a Verus pilot on one dequant/parser function (#3495) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * quorum(PMAT-3496 brief, PR #3496): 3/3 PASS on 57e0e4dfb — gemini lanes, measured; author claude-fable-5-1 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * PMAT-3577: the logit-parity receipts under contract — parity-receipt-v2, extract:parity-receipt, and the 7 back-filled Until this row the logit-parity records under evidence/parity/** had NO validator of any kind. Not a weak one — none. That is why seven of them sat in the tree carrying no comparator for months: there was nothing that could have noticed. WHAT LANDS · contracts/parity-receipt-v2.yaml — three shapes: parity-receipt-complete (closed, ignoredProperties empty), parity-comparator-self, -oracle. The subset has no sh:or, so the comparator split is two shapes over two subclasses the extractor assigns by kind. · contracts/parity-receipt-v1.yaml — the retired layout, recorded with NO shape: all instances were migrated, and a shape whose target class nothing instantiates passes vacuously. The enforcement that replaces it fires — the extractor refuses an unmigrated record BY NAME, and so does the predicate. · ontology/extract/parity_receipt.rs — record / unmigrated / other, and skipping is never silent. · scripts/parity_receipt_denominator.sh + evidence/parity/EXPECTED_RECEIPTS — the count is pinned by an INDEPENDENT predicate. An extractor checked against a number the extractor produced proves nothing. · Unknown{ExtractorMiss}, exit 2 — a new element of the verdict lattice. An extractor that silently saw the wrong corpus reports the same "no violations" as one that saw all of it. · the 7 records migrated to v2 and back-filled with comparator {kind: self, reason}, in this commit, as the row requires. THE DENOMINATOR IS 7, NOT 8. The #3574 receipt is in PR #3575, still open; it is not on main. Measured: 113 files under evidence/parity, 7 parity records, 0 with a comparator. #3575 bumps it to 8 when it lands — this row's own falsifier on its first real use. check_parity_receipt.sh IS NOT TOUCHED, and that is the finding, not an omission. It validates the THROUGHPUT family (instrument, protocol_ref, lanes[], decode_tok_per_sec, the #2696 cross-class defect); a logit record has never carried one of those keys. Folding it in — as item 6 asked — would have deleted the #2696 validator from a family nobody was watching. One validator per artifact family, and the discriminator is the artifact's required keys, never its filename. THE BACK-FILL IS A RELABEL. `raw` is the original apr parity --json document key for key; every envelope field is quoted from committed evidence named in each record's provenance.record. model_sha256 is carried only by the one record that measured it: hashing the files today and attaching that to a receipt about 2026-09-06 would be a claim about a different world wearing a witness's clothes. One derived field was wrong first time and is worth recording: result.verdict copied raw.parity — apr's own per-position flag — which says PASS for the two 1.5B cells their own RECORD.md calls RED. It now resolves the threshold from thresholds.yaml and reproduces all seven readings the RECORD.md files state, both REDs included. No threshold is ever typed into a shape. NOTHING IS ARMED. armed_shapes lives in lint-baseline.json, a shared file this row may not touch (decision 7); the shapes are computed and reported, as ladder-green was at ONT-4c1. Arming is a follow-up with the label. CONTROLS, both directions: 7 violations with the comparator stripped from all seven → 0 as committed; exactly 1 for a single plant, naming focus node and property; widening sh:in to accept `oracle` turns ont4c3_parity_receipts RED, and mutating only the real contract turns the fixture-drift test RED; an unmigrated record declines at exit 2 naming the file; 2 receipts against a denominator of 1 declines naming both numbers; pass / fail / decline are 0 / 1 / 2. cargo test -p aprender-contracts --lib ontology::extract::parity_receipt 10 ok cargo test -p aprender-contracts-cli --test ont4c3_parity_receipts 8 ok bash scripts/parity_receipt_denominator.sh --self-test 4 ok pv lint contracts --gate {sigma,relations,shapes} Pass, 0 violations pv extract contracts --check rc 0 ONT-4c3 is NOT bound in the ONT-001 ledger: that ledger is paiml/infra's paiml-ontology.md, where v4.8 still defines ONT-4c3 as KERNEL receipts. The re-scope is an infra PR, not this one — raised with the cop rather than left as a checked box. Receipt: docs/audits/impl-PMAT-3577-receipt.md. Refs #3577, #3576, #3269, #3575, #3567 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: name it WrongCorpus, not ExtractorMiss — a prefix of its opposite is a defect `ExtractorMissing` already meant the opposite thing: an extractor that does not exist. `ExtractorMiss` would have sat beside it in the same 18-element lattice, one letter apart, with the shorter a PREFIX of the longer — `grep ExtractorMiss` matches both, and any substring test over the reasons merges them silently. That is the defect class this repo keeps paying for (#3573 today), so the name now says what happened: the extractor RAN and read a corpus the tree does not declare. Renamed before it reached main, when a rename is a sed and not a migration. Caught in review by the cop. Refs #3577, #3573 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: the ONT-4c3 divergence is filed as paiml/infra#814, not left in a thread Two repos holding different definitions of one identifier is a row, not a flag in a message: nothing collides where a tool would see it, so it collides in a person's head months later when they implement the ledger's meaning and find their correct work unusable. Refs #3577, paiml/infra#814 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * #3605: removed_by gets a shape — refusal-receipt-v1, with a closed sentinel set so the required field manufactures nothing `removed_by` had 0 occurrences tree-wide, no schema and no validator. Every refusal #3597 writes would have minted an unenforced convention, and a tree full of consistent-looking `removed_by:` lines reads as validated when it is decoration. PHASE 0 — ITS OWN CONTRACT, AND NOT BECAUSE IT IS EASIER TO WRITE Option B is refused on a fact: `parity-receipt-v1` DOES NOT EXIST AT HEAD. It is in unmerged #3600, and there it is deliberately the RETIRED layout with no shape and no instances. A live validated field on a superseded contract with zero focus nodes is a field nothing can carry. Independently, by the discriminator this tree has now used twice — an artifact family is named by its REQUIRED KEYS, never by its filename — a parity receipt requires host, backend, comparator, threshold_source and per-position metrics; a refusal requires a verb, a reason, an exit code and removed_by. They share no required key. A THIRD option was considered and refused, and it is the more attractive one: put removed_by on apr-cli-commands-v1.yaml, where the verb already IS the focus node and the universe is already the trustworthy 111. Refused because the registry is the UNIVERSE and #3597's method is to DIFF the registry against the buckets. If the buckets live in the registry, the denominator and the numerator are the same artifact and the diff is vacuous by construction. The registry says what EXISTS; a refusal says what was TRIED. Keeping them apart is what lets that diff be a real diff. THE FORCED-BINDING TRAP, AND WHY minCount 1 IS SAFE HERE A required field with no escape manufactures false data: `pv validate` requires kani_harnesses so authors fabricate one, and apex#57 had lean_theorem copied verbatim into twelve contracts resolving to nothing, gate green throughout. The escape is a closed sentinel set, so "there is legitimately nothing here" is SAYABLE and still CHECKABLE: v<major>.<minor> the release that removes it never refused permanently by design unscheduled a defect with no release chosen tbd, soon, n/a, pending, a bare `0.70`, a patch-level `v0.70.1` and a git sha are all RED. A version rather than a sha because a refusal answers "which release do I need?" and a sha is precise about the tree and silent about the boundary; the sha form is made INVALID rather than discouraged, because a shape that permits two spellings gets both. BOTH ARMS, and the second is the point: 9 cases over 8 fixtures. refusal-ok carries all three declared forms and passes with 3 focus nodes; refusal-undeclared-sentinel plants the PLAUSIBLE `tbd` and must go red. Mutation proved red-capable: widening the pattern to accept tbd fails exactly a_plausible_but_undeclared_sentinel_is_refused and accept_and_refuse_are_distinct_answers, and restoring makes 9/9 green. Fixtures carry the real contract byte for byte, so widening it without them is caught too. No new extractor: entity {type: json, ref} + vocabulary is the existing machinery for "validate this document", and a second reader for one more family is the thing this tree keeps filing against. done_when 6 — the interim re-check list is EMPTY: `removed_by` still has 0 occurrences at HEAD, so no refusal written under the interim needs revisiting. The ledger is seeded with the one refusal already measured (`apr bench` refuses qwen35) so the shape has a real focus node instead of passing vacuously over an empty list. WHICH verbs land there is #3597's bucket, not this row's. Refs #3605, #3597, #3600, #3080 Pmat-Ticket: PMAT-3605 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: count parity-receipt in by_entity_type — ONT-001 v4.10's probe reads ABSENT otherwise The spec's probe asks `by_entity_type["parity-receipt"] == 7`. Measured on this branch before the change: the map carried pv-contract, gguf, apr-model, code and lean, and NO parity-receipt key. The seven records were there — by_shape showed parity-receipt-complete=7 — but the entity-type map did not carry them, so the probe would have read ABSENT. AN ABSENT KEY IS NOT ZERO. A consumer treating it as one measures nothing and calls it a pass — the same shape as #3610, one map over. Registering the entity type in Sigma was not enough; it has to be counted where the probe looks. A test now asserts the key exists AND that it equals the shape's own focus-node count, so the two numbers cannot drift apart. Refs #3577, #3610, paiml/infra#831 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3577: regenerate contracts.nt with a pv built from THIS worktree — the shared target dir served another branch's binary The merge commit's contracts.nt was 107 triples short: every parity-receipt and parity-comparator node was missing, because `cargo metadata`'s target_directory is shared across worktrees and the pv on PATH had been built from a different branch minutes earlier. AND `pv extract contracts --check` PASSED ON IT, because the check re-derives the graph with the same binary. A stale tool comparing an artifact against its own re-derivation agrees with itself about nothing being there — the derived file and the checker were wrong in the same direction, which is the only way that gate can fail to fire. Rebuilt with CARGO_TARGET_DIR pinned to this worktree; the 107 triples return and by_entity_type[parity-receipt] reads 7 rather than absent. Refs #3577 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * register the new CLI test in scripts/tree_reader_tests.txt The registry is derived from the tree and drifts the moment a test that reads the tree is added without listing it. Each of the three new tests failed this guard on its OWN branch — not a shared commit, as first read. Pmat-Ticket: PMAT-3598 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * register the new CLI test in scripts/tree_reader_tests.txt The registry is derived from the tree and drifts the moment a test that reads the tree is added without listing it. Each of the three new tests failed this guard on its OWN branch — not a shared commit, as first read. Pmat-Ticket: PMAT-3598 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * #3600: the migration broke TWO readers of the records, including the release judge I asked "who else reads this?" for the <unk> assumption and did not ask it for my own migration. Two consumers read the legacy top-level metrics: crates/apr-cli/src/commands/parity_admission.rs -> mac-check RED scripts/check_model_parity.sh -> C14, the RELEASE judge The second is the one that matters: C14 is the pre-publish dogfood's parity gate, and on the migrated records it reported "no per-position metrics in the output" for a record that is fine. The migration would have taken the release gate down. THE RULE, STATED ONCE IN BOTH READERS: a v2 receipt EMBEDS the raw `apr parity --json` document under `raw`, so look inside the envelope when there is one. A fresh `apr parity` run is the raw document itself and carries the readings at the top level. Those are two different INPUTS — a tool's output and an archived receipt quoting it — not two spellings of one, which is the distinction that makes this a rule rather than the permissiveness #3613 refuses. The self-test's fixture builder read the same way, so the must-RED twin was being built from a KeyError and that control could not have fired. Verified: 7B sentinel PASS, 1.5B sentinel RED (unchanged from pre-migration), check_model_parity.sh --self-test 26/26, parity_admission 19 tests. Refs #3600, #3577 Pmat-Ticket: PMAT-3577 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3604: the F2 hybrid guard runs once per (model sha256, apr version, device) and leaves a receipt f2_validate_qwen35 proves the CUDA hybrid forward against a 64-position CPU reference before it will serve a token, on EVERY apr run. Measured (#3598 row 1): 67 % of a 14 s time-to-first-token, 90-93 % of it the CPU forward. The guard is right to exist and wrong to run per call. Now: the first run of a (model sha256, apr version, device) triple validates and writes a receipt; a later run whose triple matches reads it and skips the forward; `apr run --revalidate` forces a fresh run and rewrites it. MEASURED HERE, RTX 4090, Qwen3.5-4B-Q4_K_M, the 144-word row, --max-tokens 1, GPU occupancy recorded before every run, binary built from this tree under a PINNED target dir (the shared one handed me another worktree's binary first): --revalidate, warm cache 17.18 s wall guard 9,747 ms on 65 positions receipt hit, warm cache 7.39 s wall guard 0 ms sha256 1,173 ms -9.8 s wall; guard 9,747 -> 0; the key costs 1.17 s/run (2.74 GB at 2.3 GB/s) and is printed separately so it cannot hide in either number. THE RECEIPT IS THE VALIDATION, WHICH IS WHY IT IS STRICT. Every path that is not "three keys match" validates, and the three planted-receipt falsifiers were run END TO END in the real binary, not only as unit tests: wrong model sha256 -> re-validated: "receipt is for model bbbb…, this file is 00fe…" wrong apr version -> re-validated: "receipt written by apr 0.61.0, this is 0.68.2" wrong device -> re-validated: "receipt written for NVIDIA GB10, this device is …4090" corrupt file -> re-validated: "receipt unreadable (…: not a receipt)" missing file -> re-validated: "no receipt for this model" Absence is never consent, and absence and unreadability are told apart. A RECEIPT IS WRITTEN ONLY AFTER A VALIDATION THAT JUDGED SOMETHING. The guard has three early exits that let the GPU serve without comparing a position — SKIP_PARITY_GATE=1, a probe under two tokens, a CPU reference that would not run. f2_validate_qwen35 now returns F2Verdict {Accepted{positions_judged}, Rejected, NotJudged} instead of bool, and only Accepted writes; otherwise a one-token prompt would "validate" the triple for every prompt after it. The decision table (f2_receipt.rs) is pure and CUDA-free, so its 13 tests run on every build. --revalidate reaches the guard through the same env seam the guard already reads SKIP_PARITY_GATE from, rather than a 38th positional parameter on run_entry::run and six forward signatures #3606 is changing. done_when 5 is partial and says so: [source=receipt|fresh] is on the guard's stderr line; the `apr run --json` field lands with #3606's StageTimings, and F2Outcome{source, validate_ms, sha256_ms, receipt_path} is returned to the call site for exactly that. #3606 and this PR both edit f2_validate_qwen35's return path; whichever lands second reconciles ~10 lines, and #3606's lane is told. Evidence: evidence/perf/3604/MEASUREMENT.md + the stderr of all nine runs. Refs #3604, #3596, #3598, #3606 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * roadmap: mint PMAT-3604 so the AD-04 quorum for #3634 can run pmat work status PMAT-3604 refused ('Item not found'): #3604 was minted as a GitHub issue with a done_when but never as a roadmap entry, and the quorum script hard-requires the work item. Fragment + aggregate, nothing else. Refs #3604, #3634 ont-delta: none — a roadmap entry; no ontology entity, shape, reason or resolution Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(run): --gpu that fell back to CPU reported success — the refusal had no caller (#3602) `apr run --gpu` on a model whose GPU attempt is rejected at runtime printed a result and exited 0. Measured on lambda (RTX 4090 sm_89, apr 0.68.2 e6f77c98c, qwen2.5-coder-0.5b-instruct-q4_k_m): exit=0 stderr: warning: GPU output diverges from CPU at position 1 (cosine 0.4153) stdout: { "used_gpu": false, "inference_time_ms": 33646.13 } 33.6 seconds on the CPU, reported as a successful --gpu run, with `used_gpu: false` as the only signal — the same value a deliberate CPU run reports. The decision was already made, recorded and unit-tested. `registry::after_generation` implements it (R-0b, #3002/#3042): a FORCED accelerator that fell to CPU is a refusal, a DEFAULT selection that fell to CPU gets a corrective line. A git grep found its only callers were its own tests. `registry::announce` and `registry::parity_line` are in the same state, which is why a real run prints zero `selected:` and zero `parity:` lines. So this is not a new policy. It is the recorded one, reached from `apr run`: - `dispatch.rs` classifies the request ONCE via `registry::Request::wanted()` rather than re-deriving "forced" — two spellings of one rule drift apart. - `run_entry::run` calls `reconcile_accelerator` before any success output. - `--json` gains a `backend` object: `requested` / `ran` / `fell_back`, because `used_gpu: false` alone collapses "ran on CPU deliberately" with "asked for the GPU and was refused it". `accel::tests::every_accelerator_surface_calls_the_refusal` passed throughout: it guards `ensure_available`, the BUILD-time refusal, not the RUNTIME one. A guard over one of two refusals reads as coverage of both. Falsifier, both directions (`run_tests_accel_reconcile.rs`, 6 cases). Planting the pre-fix behaviour (`reconcile_accelerator` → `Ok(None)`) turns `a_forced_accelerator_that_ran_on_cpu_is_refused` and `the_json_distinguishes_a_deliberate_cpu_run_from_a_rejected_gpu_run` RED while `a_forced_accelerator_that_actually_ran_on_gpu_says_nothing` stays green — so the tests discriminate rather than just failing. `used_gpu: None` is Unknown, not a fallback: refusing on it would make every non-reporting path a hard error. NOT included, deliberately: the rejection's REASON (cosine, position) is on stderr but not in the JSON. It is produced inside realizar's F2 gate and no channel carries it to the CLI; adding one is a #3606-shaped follow-up, not something to approximate with a guess here. Refs #3602, #3483 Pmat-Ticket: PMAT-3602 * test(PMAT-3346): the measured Qwen3.5-0.8B inventory the dense path cannot reproduce QE2E-INV-001 could not be judged because nothing in the tree held a MEASURED Qwen3.5 tensor inventory to judge against. This adds one: the 320 tensors of ~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782...dc517), read straight from the GGUF header rather than from a model card. It already falsifies the current arithmetic. Dense/GQA accounting applied to that file gives 644,400,128 against a measured 752,393,024 — short by 107,992,896, 14.4% of the model, because 18 of the 24 layers are Gated DeltaNet and no term here counts their conv, gate, state or output projections. Two shapes in the file are also not what dense accounting predicts, and both are pinned: attn_q is [1024, 4096] = 2 * num_heads * head_dim (the q projection emits the attention output gate alongside the query; attn_output [2048, 1024] confirms num_heads * head_dim = 2048), and attn_q_norm/attn_k_norm are present at head_dim. The file is also TIED — it has no output.weight. Refs #3346 Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(PMAT-3346): ModelConstraints carries the gated-DeltaNet shape, so a hybrid model can be counted contracts/model-families/qwen3_5.yaml declares inner_size, state_size, conv_kernel, group_count and full_attention_interval under constraints:, and ModelConstraints carried none of them. A Gated DeltaNet layer's parameters live entirely in those dimensions, so every consumer of the descriptor counted Qwen3.5 as if three quarters of its layers did not exist. Carried through as ModelConstraints::deltanet: Option<DeltaNetShape> — the runtime YAML loader (parsing.rs) and the compiled-in registry (build_parsing.rs + build_codegen.rs) both populate it, and FALSIFY-MF-QWEN35-010 pins that the declared values survive the trip and that no other family acquires a shape it never declared. qwen3_5.yaml is the only descriptor with these keys, so every other family keeps byte-identical accounting. model_arithmetic gains gated_deltanet_layer_params (one term per GGUF tensor: attn_qkv, attn_gate, ssm_conv1d, ssm_alpha/beta, ssm_a, ssm_dt.bias, ssm_norm, ssm_out) and hybrid_layers (the interleaved schedule). attention_layer_params gained two terms the real file has and dense accounting did not model: the gated q projection (2*n_h*d_k) and the q/k norm vectors. Falsified against a real model, not against itself: fed the 0.8B configuration, the equation now reproduces the 320-tensor inventory of Qwen3.5-0.8B-Q4_K_M.gguf EXACTLY — 752,393,024, both layer kinds matching tensor for tensor. QE2E-INV-001 is still NOT asserted, and no range was widened to make it pass. The 9b descriptor now yields 8,344,907,136, up from 8,208,519,168 but still 0.655B below [9.0B, 9.2B]. The remaining gap looks like descriptor drift rather than missing arithmetic: 9b keeps inner_size 2048 — the value the 0.8B uses at hidden_dim 1024 — while quadrupling hidden_dim, and its group_count 8 fails 8 * 128 == 2048, a consistency the measured 0.8B satisfies at 16 * 128. Only a real Qwen3.5-9B file can settle it; none is on this box. Refs #3346, #3347 Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * contracts(binding): the QE2E-INV-001 note blamed a gap that is now closed The note said the obligation was undischarged because ModelConstraints does not carry the DeltaNet shape keys. It does now, and the 0.8B count reproduces the real GGUF exactly. What actually blocks the obligation is descriptor drift at the 9b variant. Pmat-Ticket: PMAT-3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3346 (adoption): the QE2E-INV-001 note lost its closing quote, and pv extract shrank the graph by 356 triples instead of refusing d3cc76f1f rewrote the note on the QE2E-INV-001 binding and dropped the trailing `"` — contracts/binding.yaml stopped being valid YAML at line 764 (`found unexpected end of stream`). Nothing in the PR noticed because `pv extract contracts` does not refuse a binding registry that will not parse: it emitted a graph with 15,244 triples where main has 15,600 — every bound symbol AFTER the broken entry (prune::run, distill::run, harness_ir::*, ptx_explain::run, …) silently gone — and `--check` would have agreed with itself. Found while regenerating the derivative for this adoption, by the drop, not by any gate. One character. With it, binding.yaml parses (156 entries, same as main) and the extraction is byte-identical to main's committed contracts.nt, so this PR owes no graph change after all. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(roadmap): mint PMAT-3602 as a work item so the AD-04 quorum can resolve it `quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work` reads the aggregated roadmap rather than an id counter. The branch carried the trailer `Pmat-Ticket: PMAT-3602` while no such work item existed — a GitHub issue number is not a work item. The id is DERIVED from the issue (`pmat work add --github-issue` calls that "the path to prefer": GitHub allocates centrally, so two agents cannot be handed the same id, which `max(id)+1` cannot promise). Fragment, not a direct edit to roadmap.yaml: additive, one file, and it does not make every stacked branch dirty on the same lines. Acceptance is hand-entered from the issue's done_when — `pmat work add` derives neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414). Guards: diff_additive, fragment_required, ids_unique, sorted, completion_is_cited all exit 0; `make roadmap-aggregate-check` reports `roadmap.yaml == aggregate(48 fragment(s)), idempotent`. Refs #3602 Pmat-Ticket: PMAT-3602 * roadmap: PMAT-3637 — the registration row this PR is, so its quorum judges fidelity not implementation Round 0 ran with --ticket PMAT-3604 and all three lanes FAILed for the same reason: the diff implements none of PMAT-3604's criteria. Correct — it was never meant to; it registers the ticket so #3634's quorum can start. The receipt is kept as quorum-PMAT-3637-r0-misframed-as-3604.json (a FAIL is a record, not a mistake to erase). PMAT-3637 is the row this diff satisfies: fidelity of the PMAT-3604 entry to issue #3604's done_when, additive aggregate, nothing outside docs/roadmaps/. Round 1 runs against it. Refs #3604, #3634 ont-delta: none — roadmap entries only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(roadmap): mint PMAT-3605 as a work item so the AD-04 quorum can resolve it `quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work` reads the aggregated roadmap rather than an id counter. This branch carried the trailer while no such work item existed — a GitHub issue number is not a work item until something mints it. The id is DERIVED from the issue: `pmat work add --github-issue` calls that "the path to prefer", because GitHub allocates centrally and `max(id)+1` cannot promise two agents different ids. A fragment rather than a direct roadmap.yaml edit — additive, one file, and it does not make every stacked branch dirty on the same lines. Acceptance hand-entered from the issue's done_when; `pmat work add` derives neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414). Guards: diff_additive, fragment_required, ids_unique, sorted, completion_is_cited all exit 0; `make roadmap-aggregate-check` idempotent. Refs #3605 Pmat-Ticket: PMAT-3605 * PMAT-3346: the roadmap row, so the AD-04 quorum for #3350 has a ticket to judge against Acceptance transcribed from issue #3346 as this PR answers it (the type that reads the descriptor was wrong, not the range or the descriptor), with the 9B range instantiation explicitly out of scope until a real 9B GGUF exists. Refs #3346 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * roadmap: round 1 said two true things — fix both Lanes 1+2 (gemini-3.1-pro-high, gemini-3.8-flash-high) FAILed PMAT-3637 on (a) the round-0 receipt committed under docs/audits/, which criterion (3) forbade — dropped; the receipt judged the wrong question and its ticket field would read as a verdict on PMAT-3604; and (b) a criterion in the PMAT-3604 notes that issue #3604 does not state ('Only an Accepted verdict writes a receipt') — it is PR #3634's design decision, not a done_when item; removed from the transcription. Criterion (3) now admits this row's own receipt at docs/audits/quorum-PMAT-3637*.json, which it must, or the PASS receipt could never be committed. Refs #3604, #3634 ont-delta: none — roadmap entries only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(run): quorum round 1 found the --json surface skipped and a branch that could never run Two findings from lane 1 (gemini-3.1-pro-high) of the AD-04 quorum, both correct, both fixed. Lane 3 passed the PR without seeing either. 1. `reconcile_accelerator(...)?` ran BEFORE `print_run_output`, so a rejected `--gpu` run early-returned and `--json` emitted nothing at all — the exact surface #3602 item 1 names. Machine surfaces (`--json`, `--stream`) now emit before the refusal propagates; the human surface still prints nothing. This is a DELIBERATE deviation from `after_generation`'s contract, which says the caller "must print NO output" on a forced refusal. Named rather than quiet: that rule exists so a CPU result is never read as a GPU success, and a document carrying `"backend": {"fell_back": true}` beside exit 14 cannot be read that way, while a human-formatted success blob can. The protective half is kept; the half that blinded `--json` consumers is not. 2. The caller's `if let Some(note)` arm was UNREACHABLE. `announced` was `Some("gpu")` exactly when `forced` was true, and `after_generation`'s corrective-line branch requires `forced == false` — so the Some arm could never execute. A branch with no reachable caller, which is precisely the defect this PR fixes, reproduced one layer down while fixing it. `reconcile_accelerator` now returns `Result<()>`, wires the FORCED half only, and says in its doc comment why the default-selection half is not wired: nothing calls `registry::announce`, so there is no recorded announcement for a default selection, and manufacturing one would assert a choice this process never made. That is REG-8, not this PR. Verification: cargo test -p apr-cli --lib = 7289 passed, 0 failed, 12 ignored; cargo fmt --all --check = 0; cargo clippy -p apr-cli --lib -D warnings = 0. Refs #3602 Pmat-Ticket: PMAT-3602 * PMAT-3346 (adoption): a bias vector is as wide as the projection it biases Found by the AD-04 quorum on #3350 (lane 1, gemini-3.1-pro-high, cited model_arithmetic.rs:144): the projection term used q_out for a gated family's q matrix (2*n_h*d_k — attn_q emits the output gate, MEASURED in Qwen3.5-0.8B) while the bias term still used q_dim. A bias narrower than its projection is not a model. One token: q_dim -> q_out in the bias sum. Why a delta-0 measurement did not catch it: no shipped family exercises the case. Qwen3.5 has no attention bias; Qwen2.5 has biases but is not gated, so q_out == q_dim there. The test that pinned 12 (a q_dim bias under a q_out matrix) now asserts 16 and says why, and a second test holds the other polarity — a non-gated family with biases is unchanged at 12. 115 model_arithmetic + model_family tests pass; oracle 218; clippy clean. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(run): quorum round 2 — a --json --benchmark leak, a comment claiming coverage it lacks, and a stray artifact I swept in Three findings, two lanes, all correct. 1. LANE 1 — `--json --benchmark` leaked a human success blob on a refused run. The refusal path carried its OWN copy of "is this a machine surface" (`stream || output_format == "json"`) and dropped `!benchmark`. So that combination entered the branch, matched neither machine arm inside `print_run_output`, and fell through to the human benchmark rendering — for a run being refused. Two spellings of one condition drifting apart in the gap between them, which is the defect this PR's own dispatch comment warns about. One spelling now: `emits_machine_output(stream, output_format, benchmark)`. `the_machine_output_predicate_matches_print_run_output` pins it to the arms it describes over every flag combination; planting the pre-fix spelling turns it and `json_plus_benchmark_is_not_a_machine_surface` RED, verified. 2. LANE 1 — the classifier comment claimed `--gpu-layers all|n` was among the flags handled here. It is not: `Commands::Run` carries `gpu` and `no_gpu` and nothing else; `--gpu-layers` belongs to `apr serve`. So `layers_want_accelerator: false` is CORRECT and the comment was the defect — a comment asserting coverage the code does not have. Both comments now say the input is genuinely absent for this surface rather than stubbed. 3. LANE 3 — `trace-1789944278.json` (232 lines) was committed to the repo root. A runtime byproduct of `test_print_chrome_trace_creates_file`, which writes `trace-<epoch>.json` into the CWD when given no `--trace-output`. I swept it in with `git add -A` immediately after a quorum round — the exact thing my own notes say never to do there. Removed, and this commit stages files by name. The test writing into the repo root is a separate defect and is not fixed here. Round 2 was 2 FAIL / 1 NO-VERDICT. Lane 2 has now returned NO-VERDICT twice with `envelope_status: SUCCESS`, `transport_status: SUCCESS` and non-empty `raw_bytes` — it answers, and the harness cannot extract a verdict from what it returns, so width 3 has been delivering two votes. Verification: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored; cargo fmt --all --check = 0. Refs #3602 Pmat-Ticket: PMAT-3602 * roadmap: transcribe #3604 verbatim — round 2's pro lane was right Criterion (5) had grown an implementation-status clause from PR #3634's report ('the --json field lands with #3606 StageTimings…') that issue #3604 does not state, and the out-of-scope list was a paraphrase. PMAT-3637 (1) says every criterion present, none added: the notes now carry the issue's done_when 1-6, admission rule and out-of-scope list as written. Refs #3604, #3634 ont-delta: none — roadmap entries only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * audit: quorum-PMAT-3637 — 3/3 PASS, round 3 on the ruled trio pro / 3.7-flash / 3.6-flash, three conversations, no silences. Rounds 0-2 each said something true about this PR: round 0 judged the wrong ticket; round 1 found a receipt outside scope and an invented criterion; round 2 found an implementation-status clause in a 'none added' transcription. None is committed — each judged a diff that no longer exists. Refs #3604, #3634 ont-delta: none — audit artifact only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3346: quorum verdict 3/3 on e060c7e2c (AD-04) Round 0 was 1 FAIL / 1 no-verdict / 1 PASS and the FAIL was real (the bias width, fixed in e060c7e2c). Round 1 on the fixed head: 3/3 PASS, gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent. Refs #3346, #3350 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(audits): AD-04 quorum receipt for PMAT-3577 — AGREED 3/3 Three independent agy lanes on distinct models, none in the author's family (author measured as Opus 5 / claude): gemini-3.1-pro-high = PASS gemini-3.7-flash-high = PASS gemini-3.6-flash-high = PASS Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG rather than by editing the shared config, which three sessions were launching against concurrently. THE MERGE DID NOT CHANGE WHAT WAS JUDGED. The lanes ran against 7f82e2cd3; the PR head is 4fbd2ec31, a merge of main. The trees differ by four files, but all four arrived FROM main, and the PR's own contribution relative to main is byte-identical across the merge: files, pre-merge : 49 files, post-merge : 49 only in post : (none) only in pre : (none) sha256(diff vs merge-base), pre-merge : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35 sha256(diff vs merge-base), post-merge : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35 So this is the same head for review purposes and does not need a new round. A naive `git diff origin/main..HEAD | sha256sum` from my local head disagrees only because that head carries this receipt commit, which the lanes never saw — comparing the receipt's parent is what makes the two comparable. Refs #3600 Pmat-Ticket: PMAT-3577 * fix(tests): the shape-count ratchet was red — #3605 adds a 6th shape (quorum round 1) `the_tracked_repo_graph_is_fresh` asserts `shapes_n == 5` and names the five. `refusal-receipt-v1` makes six, so this PR shipped a CI red that a quorum lane found and I did not. MEASURED, not guessed: `pv extract contracts --check` on this branch reports `shapes_n: 6, triples: 15620`. The hardcoded count stays hardcoded. A shape added without anyone noticing is exactly what this assertion exists to prevent, so adding one is SUPPOSED to turn it red and make you name the new shape. Noted in the comment that the counter is shared across branches — a sibling PR adding a shape (#3600's parity-receipt-v2) will need it raised again at merge, which is the ratchet working rather than a conflict to route around. Refs #3605 Pmat-Ticket: PMAT-3605 * chore(audits): AD-04 quorum receipt for PMAT-3602 — AGREED 3/3 gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS. Author measured as Opus 5 (claude); no lane in the author's family. Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared config that other sessions were launching against. This trio was chosen on measured odds after flash-class lanes returned NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my earlier runs). All 15 lanes voted in this batch. Refs #3638 Pmat-Ticket: PMAT-3602 * chore(audits): AD-04 quorum receipt for PMAT-3605 — AGREED 3/3 gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS. Author measured as Opus 5 (claude); no lane in the author's family. Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared config that other sessions were launching against. This trio was chosen on measured odds after flash-class lanes returned NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my earlier runs). All 15 lanes voted in this batch. Refs #3613 Pmat-Ticket: PMAT-3605 * PMAT-3351 (adoption): re-id the fragment — PMAT-3347 is #3348's row on main Both #3348 (merged; bound the qwen35-e2e equations, closed #3347) and this PR (the L2 column reads a link, not an index) address issue #3347, and both claimed roadmap id PMAT-3347. Two PRs cannot share a row: main's PMAT-3347 stays as #3348's, this PR's fragment becomes PMAT-3351 (github_issue 3347, same title, same notes), and the aggregate is rebuilt from main's roadmap.yaml plus this branch's fragments — 935 + 1 = 936, nothing re-serialised. Refs #3347, #3348 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3351: notes transcribe issue #3347's Ask verbatim (bullets 2 and 3; bullet 1 was #3348) The quorum lanes read pmat work status, i.e. this fragment's notes, not the GitHub issue. Empty notes would have them judge against a title. The issue has no done_when section, so its Ask and its two defect statements are transcribed as written, with the out-of-scope bullet named and no count offered as a criterion. Refs #3347 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3338: register the second ticket this PR closes, so the quorum judges the truncate fix as in scope Quorum round 0 on a4b76397f was 2 FAIL / 1 PASS, both FAILs on one point: the branch's first commit fixes #3338 (--table panicked on a byte-index cut inside a multi-byte char) and PMAT-3351 does not ask for it. Both lanes called the in-scope work correct. The fix is a prerequisite — the L2 column is exercised through --table on the real corpus, which panicked before it — and #3338 is an OPEN issue this PR genuinely closes. So it gets its row, transcribed from the issue, and round 1 runs with --ticket PMAT-3351,PMAT-3338. Refs #3338, #3347 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3604: the receipt's temp file is private to its writer — a shared name made the atomicity claim false AD-04 quorum round 0 on #3634, lane 1 (gemini-3.1-pro-high), cited f2_receipt.rs:180: write_receipt used one shared <sha>.json.tmp, so two apr runs validating the same model at once could have writer B truncate the file writer A was about to rename, and A rename B's partial into place. The reader would classify it Unreadable and validate — the safe direction, done_when 4 — but the docstring said 'never a truncated one', and that was false. The temp name now carries the pid and a per-process counter; a failed rename removes its own temp. A racing test spawns two writers on one path forty times and parses the survivor each time: it is always one writer's WHOLE receipt. 14 tests. Lane 1's other two findings, for the record: --revalidate is done_when 2 verbatim (the fragment carries it), not scope creep; and F2Outcome's fields are the values the stderr line already prints and done_when 5's data path for #3606's JSON, not dead code. Two lanes passed the same diff. Refs #3604 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3604: quorum verdict 3/3 on dd546592b (AD-04) Round 0 was 2 PASS / 1 FAIL; the FAIL was the shared temp-file race, fixed in dd546592b. Round 1 on the fixed head: 3/3 PASS, gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent. The PMAT-3604 roadmap row was on this checkout UNCOMMITTED for the resolver (pmat work status reads the checkout); it lands via #3637. The lanes judged the committed diff origin/main...HEAD. Refs #3604, #3637 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * PMAT-3351 + PMAT-3338: quorum verdict 3/3 on d915b8f16 (AD-04) Round 0 (--ticket PMAT-3351 alone) was 2 FAIL / 1 PASS, both FAILs on the #3338 truncate fix being out of scope; both called the in-scope work correct. Round 1 with both tickets registered and named: 3/3 PASS, gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent. Refs #3347, #3338 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(contracts): the v1 contract had no equations, and its own denominator guard was never wired Two reds on this PR, both mine, and the second is the better finding. 1. `contract_data_integrity` is a SHRINK-ONLY ratchet at 444 and this branch made it 445. The 445th is `parity-receipt-v1: no equations`. Measured which one rather than guessed — `parity-receipt-v2` is not in the list. `kind: pattern` correctly declares no SHAPE (a shape over a class nothing instantiates passes vacuously, which is #3610's whole lesson), but equations are not shapes, and this contract does assert things: PRC1-INV-001 and PRC1-INV-002 already said them. They are now written as the two equations the integrity check reads — `no_instances` and `refused_by_name` — rather than as new claims invented to satisfy a counter. 2. `check_guards_are_wired.sh`: `NEW: parity_receipt_denominator.sh`. I ADDED A GUARD IN THIS PR AND WIRED IT INTO NOTHING — the second, independent reader of the receipt corpus, shipped where no workflow names it. Minutes before finding this I wrote into that same contract's equations, as a precondition: "both readers are wired into CI — a refusal nothing runs is not a refusal." My own precondition, violated by my own PR, in the same file. `check_guards_are_wired.sh` caught what I had just finished writing down. Now named in `ci.yml` beside the other cargo-free, model-free case tables. Its 4-case self-test passes: a legacy record refused by name, a receipt added without bumping the denominator disagreeing, bumping it making them agree. NOTE — THIS TOUCHES `.github/workflows/ci.yml`, which CLAUDE.md lists as a check-in item. Additive only: one step in an existing guard block, no matrix, trigger or gate logic changed. Wiring an unwired guard is the minimum the failing check asks for, and leaving it unwired to avoid the file would be keeping a refusal that never runs. VERIFICATION cargo test -p aprender-contracts --test validate_contracts contract_data_integrity 1 passed bash scripts/parity_receipt_denominator.sh --self-test rc 0, 4/4 bash scripts/check_guards_are_wired.sh PASS (4 -> 3) pv validate contracts/parity-receipt-v1.yaml 0 errors, valid DIFF MOVEMENT, for the receipt: the judged diff DID change — an equations block and one CI step. No behaviour the lanes reviewed was altered; the extractor, the shapes and the denominator script are byte-identical. AD-04 re-review is the inspector's call. Refs #3577 Pmat-Ticket: PMAT-3577 * PMAT-3401 (adoption): the three derivatives 24 new contracts oblige — census, graph, README — regenerated together This PR adds 24 contracts and regenerated none of the tracked artifacts derived from the corpus. On #3581 that omission surfaced one per CI round, each masked by the one before it. All three here, at once, with a pv built from this tree under a pinned target dir: contracts/census.json 1800 -> 1824 (+24, the contracts added) contracts/contracts.nt 15,600 -> 15,696 triples (+96 = 24 x 4; GREW — a drop is the tell for a malformed input or a stale binary; binding.yaml still parses, 156 entries) README CONTRACT_COUNT 2 blocks -> 1824 via make readme-sync The merge took main's generated README blocks over the branch's hand-typed 1866, then readme-sync wrote the measured 1824; the branch's number described a tree that never existed on main. Verified: all 24 contracts pv-validate under the pinned binary (control: main's crux-A-01 valid under the same one); lint_passes_on_real_contracts green, so the sigma prose ratchet holds; 1666 engine tests; ont4b shapes gate 11/11; test-binding ratchet; readme-sync-check; FALSIFY-README-002; roadmap additive added=26 deleted=0, aggregate idempotent, 26 fragments present. Refs #3401 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(run): extract reconcile_and_emit — run() hit cognitive 27 against a 25 ceiling `check_complexity_ratchet.sh` went RED with: RED NEW crates/apr-cli/src/commands/run_entry.rs::run cyclomatic 21 cognitive 27 (over a threshold, absent from the comparand) The reconcile-then-emit block I added in this PR is what took `run` over. Moved into `reconcile_and_emit`, which also gives the ordering decision — machine surfaces emit on a refusal, the human surface does not — a place to be documented that is not the middle of a 30-argument entry point. Now: PASS (D2): e6f77c98c vs 19750b384 measured by pmat 3.41.1 — none new, none grown. NOTE FOR THE NEXT PERSON: I first re-ran the ratchet against my WORKING TREE and saw the identical cognitive 27, and briefly concluded the extraction had not helped. It had. The script measures two REVISIONS — it prints `merge HEAD <sha>` — so an uncommitted fix is invisible to it. Commit, then measure. Tests: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored; fmt --check = 0; clippy -D warnings clean. Refs #3602 Pmat-Ticket: PMAT-3602 * PMAT-3496: a registration row for the registration PR — the author's earlier receipt named this ticket but no row existed on the branch * PMAT-3496: clause (1) transcribes issue #3495's title verbatim (132 Kani harnesses; Verus pilot), not a paraphrase * fix(guard): a step NAME read as a bare guard_tree.sh dispatch hid four dark guards; the #3305 guard runs nowhere and dies on intel's python 3.10 (#3644) Refs #3644 #3646 #3305 #3626 THE FINDING, proven by mutation before it was argued. check_guards_are_wired.sh counts a guard wired-by-dispatch when a workflow runs guard_tree.sh in a mode whose --dry-run RUN set contains it. Its invocation regex accepts `guard_tree.sh` followed by `'` (for a quoted run: string). ci.yml has two step NAMES, `"guard_tree.sh's own case table …"` and `"guard_tree.sh's parallel dispatcher …"`; the apostrophe matched, neither line says --no-cargo, so each was read as a bare dispatch in mode "all" -- 55 cargo-classified guards counted wired that `guard-tree` (--no-cargo) never runs and `guard-cargo` never names. Rewriting those two names (nothing else) turns the meta-guard RED: 3 -> 7 unwired -- check_model_ladder.sh, check_pathonly_devdeps_unused_in_src.sh, check_pr_review_counts.sh, check_pr_review_receipt.sh. Three of the four are cargo-classified by a COMMENT that says the build tool's name; none of those three invokes it. check_pathonly_devdeps_unused_in_src.sh (#3305: src/ may not use a dev-dep that publishing deletes; clean-room red 8/8 across two releases) has therefore run nowhere since it was written on 2026-09-15. And it could not have run on the intel hosts: `import os, re, sys, tomllib` at module level, python 3.10.12 there with neither tomllib nor tomli (measured on mac-server tonight; the sovereign-ci container has no python3 at all). Reproduced with python3.10 -S: the traceback is swallowed by `out=$(scan … || true)` and the case table prints "FAIL: missed hit_use" ×4 -- the regex reported broken by an interpreter that never ran it. Same class as #3626's guard on intel-clean-room-6. WHAT CHANGES scripts/check_guards_are_wired.sh - `_not_a_name_line` drops `name:` lines before the invocation match, in dispatcher_wired() and the per-guard scan. A step name is documentation whatever it contains. - rows 8-10: the exact ci.yml fixture (`- name: "guard_tree.sh's own case table …"`, no run: line) must leave the dispatched guard unwired; a QUOTED run: line still dispatches (the control the `'` exists for); a step named after a guard wires nothing. Mutant (drop removed): rows 8 and 10 RED. - the ledger's ratchet is now set-aperture, owned by this file. scripts/lib_baseline_ratchet.sh, scripts/check_baseline_ratchets.sh - set-aperture gains a NAME-entry admission for sets of FILES: an added entry with no `:` is admitted iff the comparand carries that file (at the entry's path, or beside the owning guard -- the ledger names siblings by basename and guard_tree.sh reads it so) AND the owning guard changed in the diff. A file this branch created is refused (PERF-028's shape). Six rows: predates -> admitted; branch WROTE it -> refused; no guard edit -> refused; escaping path -> refused; beside the OWNER -> admitted; beside a different owner -> refused. Lib mutants (always admit / branch removed / sibling resolution removed) each turn rows RED; a separate absolute-path check survived its mutant because `git cat-file -e` already refuses those, so it is not there. - classify: unwired_guards_baseline.txt -> set-aperture, owner named. scripts/unwired_guards_baseline.txt 3 -> 6, as APERTURE REVEALS with a reason per line (each may only leave): check_model_ladder.sh (T-2 release gate, the multiplatform_dogfood class); check_pr_review_counts.sh (RED on main today, "6 row(s) disagree" -- #3646); check_pr_review_receipt.sh (takes a receipt path; nothing passes one -- #3646). check_pathonly_devdeps_unused_in_src.sh is NOT ledgered: see below. scripts/check_pathonly_devdeps_unused_in_src.sh - readers tomllib -> tomli -> ENV rc=2 naming python version and $RUNNER_NAME. No purpose-built manifest reader: a second TOML implementation over 78 manifests, where a wrong "has source" is a silent PASS on exactly the #3305 class, is not a fallback worth shipping. UNMEASURED-with-a-name is the honest state on a runner with no TOML library. - the scanner ends with `SCAN-DONE manifests=N reader=…`; the shell REQUIRES it. Output without it (no reader, /bin/false, no python3) is ENV rc=2 -- never rc=1, never "no violations". The selftest and the tree scan propagate 2 instead of swallowing it. - five death rows: both readers absent (sys.modules nulled so the imports fail for real); interpreter exits 1 silently; no interpreter (127); the working interpreter as control; and THE WHOLE GUARD under the intel shape (nested, recursion-guarded) -- the call site is where `|| true` hid it. Mutants: trailer not required (3 rows RED); trailer printed on the no-reader path (1 RED); `|| true` restored (the whole-guard row RED). - three comments and one FAIL string reworded so the file no longer says the build tool's name with a space after it: it invokes none, and that substring is what CARGO_RE classifies on. guard_tree.sh --dry-run --no-cargo now lists it as `run:`; bare run on this tree rc=0 in 3 s. Passes: check_guards_are_wired --self-test 10/10; check_baseline_ratchets --self-test 61 rows; pathonly --selftest under py…
#3346 asked which was wrong, the range or the descriptor. The answer is neither — it was the type that reads the descriptor. Stacked on #3348.
contracts/model-families/qwen3_5.yamldeclaresinner_size,state_size,conv_kernel,group_countandfull_attention_interval.ModelConstraintscarried none of them, so every consumer counted a hybrid model with dense/GQA accounting and silently missed the conv, the gates, the state norm and the mixer projections of 18 of every 24 layers.Ground truth first, and it now matches exactly
The measurement is against a real file, not against the formula's own assumptions:
~/models/Qwen3.5-0.8B-Q4_K_M.gguf(sha256bd258782…dc517), GGUF header parsed directly — 320 tensors, 752,393,024 parameters.Two shapes in that file contradict dense accounting and are now modelled, both read off the tensors rather than assumed:
attn_qis[1024, 4096]= 2·n_h·d_k — the q projection emits the output gate, andattn_output [2048,1024]proves only q is doubled.output.weight.The mutation: drop the
attn_gateterm andqwen35_0_8b_config_derived_count_equals_the_measured_gguf_inventorygoes RED, short by exactlyd·inner_size= 2,097,152.QE2E-INV-001 is still NOT asserted, and the range was not widened
P(9B) moves 8,208,519,168 → 8,344,907,136, still 0.655 B below [9.0B, 9.2B]. The residual is not arithmetic — the 9b descriptor contradicts itself: it keeps
inner_size: 2048(the value the 0.8B uses athidden_dim1024) while quadruplinghidden_dim, and itsgroup_count: 8failsgroup_count · state_size == inner_size(8·128 ≠ 2048), a consistency the measured 0.8B satisfies at 16·128.Only a real Qwen3.5-9B file settles that, and none is on this box, so the obligation stays unproved and the number is pinned by a test instead of asserted.
DeltaNetShape::heads_span_the_mixer()now catches the descriptor's self-inconsistency directly.Two disclosures
deltanet: None,fixture literals inapr-clitests/oracles. Zero logic change. The enum-payload alternative breaks the same crate, becauseapr-climatchesAttentionType::HybridGatedDeltaNetin two files.qk_normis left alone deliberately. The descriptor says the family has none, the real file hasattn_q_norm/attn_k_normon its attention layers, andtensor_expectation.rsasserts the opposite. Both are true of different layer kinds — the repo should settle whetherqk_normis per-family or per-layer-kind. It is worth 4,096 parameters at 9B, so it does not touch the verdict.The roadmap write
pmat work addproduced was reverted: #3297 moved entries to fragments, and a monolithic write is exactly the contention that PR removed.Gate:
aprender-core --lib model_arithmetic28 passed,aprender-contracts --lib1526 passed, clippy-D warningsclean, fmt clean.Refs #3346, #3347, #3091, #3114
no-close: #3346 stays open — which side of the 9b descriptor is wrong is undecided until a real Qwen3.5-9B GGUF is measured.
ont-delta: shape ModelConstraints gains inner_size, state_size, conv_kernel, group_count and full_attention_interval, and model_parameter_count gains the gated-DeltaNet layer term; no obligation is added or removed, and QE2E-INV-001 is not discharged.
🤖 Generated with Claude Code