Skip to content

fix(format): ModelConstraints dropped the gated-DeltaNet shape, so 18 of every 24 Qwen3.5 layers were counted as if their tensors did not exist - #3350

Closed
noahgift wants to merge 8 commits into
mainfrom
PMAT-3346-gdn-shape-keys
Closed

noahgift wants to merge 8 commits into
mainfrom
PMAT-3346-gdn-shape-keys

Conversation

@noahgift

Copy link
Copy Markdown
Contributor

#3346 asked which was wrong, the range or the descriptor. The answer is neither — it was the type that reads the descriptor. Stacked on #3348.

contracts/model-families/qwen3_5.yaml declares inner_size, state_size, conv_kernel, group_count and full_attention_interval. ModelConstraints carried none of them, so every consumer counted a hybrid model with dense/GQA accounting and silently missed the conv, the gates, the state norm and the mixer projections of 18 of every 24 layers.

Ground truth first, and it now matches exactly

The measurement is against a real file, not against the formula's own assumptions: ~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782…dc517), GGUF header parsed directly — 320 tensors, 752,393,024 parameters.

measured GGUF tensor sum   752,393,024
model_parameter_count      752,393,024      delta 0
  gated-DeltaNet layer      21,555,360  ==  measured blk.0
  full-attention layer      18,352,640  ==  measured

Two shapes in that file contradict dense accounting and are now modelled, both read off the tensors rather than assumed:

  • attn_q is [1024, 4096] = 2·n_h·d_k — the q projection emits the output gate, and attn_output [2048,1024] proves only q is doubled.
  • the file is tied: it has no output.weight.

The mutation: drop the attn_gate term and qwen35_0_8b_config_derived_count_equals_the_measured_gguf_inventory goes RED, short by exactly d·inner_size = 2,097,152.

QE2E-INV-001 is still NOT asserted, and the range was not widened

P(9B) moves 8,208,519,168 → 8,344,907,136, still 0.655 B below [9.0B, 9.2B]. The residual is not arithmetic — the 9b descriptor contradicts itself: it keeps inner_size: 2048 (the value the 0.8B uses at hidden_dim 1024) while quadrupling hidden_dim, and its group_count: 8 fails group_count · state_size == inner_size (8·128 ≠ 2048), a consistency the measured 0.8B satisfies at 16·128.

Only a real Qwen3.5-9B file settles that, and none is on this box, so the obligation stays unproved and the number is pinned by a test instead of asserted. DeltaNetShape::heads_span_the_mixer() now catches the descriptor's self-inconsistency directly.

Two disclosures

  • Outside scope, unavoidable: five one-line deltanet: None, fixture literals in apr-cli tests/oracles. Zero logic change. The enum-payload alternative breaks the same crate, because apr-cli matches AttentionType::HybridGatedDeltaNet in two files.
  • qk_norm is left alone deliberately. The descriptor says the family has none, the real file has attn_q_norm/attn_k_norm on its attention layers, and tensor_expectation.rs asserts the opposite. Both are true of different layer kinds — the repo should settle whether qk_norm is per-family or per-layer-kind. It is worth 4,096 parameters at 9B, so it does not touch the verdict.

The roadmap write pmat work add produced was reverted: #3297 moved entries to fragments, and a monolithic write is exactly the contention that PR removed.

Gate: aprender-core --lib model_arithmetic 28 passed, aprender-contracts --lib 1526 passed, clippy -D warnings clean, fmt clean.

Refs #3346, #3347, #3091, #3114

no-close: #3346 stays open — which side of the 9b descriptor is wrong is undecided until a real Qwen3.5-9B GGUF is measured.

ont-delta: shape ModelConstraints gains inner_size, state_size, conv_kernel, group_count and full_attention_interval, and model_parameter_count gains the gated-DeltaNet layer term; no obligation is added or removed, and QE2E-INV-001 is not discharged.

🤖 Generated with Claude Code

@noahgift noahgift added this to the 0.68.0 milestone Sep 16, 2026
@noahgift
noahgift enabled auto-merge September 16, 2026 07:48
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

§13.11 rung 1 — quorum shadow verdict

S13-SHADOW pr=3350 head=abf8b56acd24be766f3eb806d08216e321926715 verdict=REFUSE class=Q1 arm_rc=1

Shadow mode: this records a verdict and merges nothing. A refusal
to arm is not a block (§13 adds zero rows to §7) — the pull request is
exactly as green as it was.

@noahgift noahgift modified the milestones: 0.68.0, 0.70.0 Sep 16, 2026
@noahgift
noahgift added this pull request to the merge queue Sep 16, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to a conflict with the base branch Sep 16, 2026
@noahgift noahgift modified the milestones: 0.70.0, 0.71.0 Sep 17, 2026
@noahgift

Copy link
Copy Markdown
Contributor Author

stale — APR-RELEASE-001 §6, applied mechanically 2026-09-20.

No activity since 2026-09-17 13:44Z; 2 release trains have been cut since (v0.68.2, v0.68.1). Age is the only input to this sweep — it is not a judgement on the work, and a train label such as pp-066/carry-* is not an exemption.

What happens next: if one more train passes while this is still labelled stale, §6 closes it. The branch is kept — closing a PR here discards a review thread, never commits.

To clear it: push, rebase, or say on the thread what it is waiting for. Any of the three removes the label at the next sweep. If it is blocked on something external, name that here — a blocker with an owner is not sprawl, and it stops the clock.

@noahgift

Copy link
Copy Markdown
Contributor Author

Adopting this PR (2026-09-21 sweep; author session is not live — branch cut 09-16, no session on the box predates it). Plan, before touching anything: the DIRTY state is the squash-of-the-stack-base class — c5b7bce70 on this branch is byte-identical (git patch-id) to main's dfd3d5f96 (#3348's squash), so every hunk built on it conflicts. Fix is git rebase --onto origin/main c5b7bce70: replay only the three commits that are this PR's own, no content change, then run the model_family / model_arithmetic tests and the roadmap guards on the rebased head before a --force-with-lease push. If a real conflict appears in the replay I stop and report it here rather than resolve it by guessing. Refs #3346.

noahgift and others added 4 commits September 21, 2026 00:34
…annot reproduce

QE2E-INV-001 could not be judged because nothing in the tree held a MEASURED
Qwen3.5 tensor inventory to judge against. This adds one: the 320 tensors of
~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782...dc517), read straight from
the GGUF header rather than from a model card.

It already falsifies the current arithmetic. Dense/GQA accounting applied to
that file gives 644,400,128 against a measured 752,393,024 — short by
107,992,896, 14.4% of the model, because 18 of the 24 layers are Gated DeltaNet
and no term here counts their conv, gate, state or output projections.

Two shapes in the file are also not what dense accounting predicts, and both are
pinned: attn_q is [1024, 4096] = 2 * num_heads * head_dim (the q projection
emits the attention output gate alongside the query; attn_output [2048, 1024]
confirms num_heads * head_dim = 2048), and attn_q_norm/attn_k_norm are present
at head_dim. The file is also TIED — it has no output.weight.

Refs #3346

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… a hybrid model can be counted

contracts/model-families/qwen3_5.yaml declares inner_size, state_size,
conv_kernel, group_count and full_attention_interval under constraints:, and
ModelConstraints carried none of them. A Gated DeltaNet layer's parameters live
entirely in those dimensions, so every consumer of the descriptor counted
Qwen3.5 as if three quarters of its layers did not exist.

Carried through as ModelConstraints::deltanet: Option<DeltaNetShape> — the
runtime YAML loader (parsing.rs) and the compiled-in registry (build_parsing.rs
+ build_codegen.rs) both populate it, and FALSIFY-MF-QWEN35-010 pins that the
declared values survive the trip and that no other family acquires a shape it
never declared. qwen3_5.yaml is the only descriptor with these keys, so every
other family keeps byte-identical accounting.

model_arithmetic gains gated_deltanet_layer_params (one term per GGUF tensor:
attn_qkv, attn_gate, ssm_conv1d, ssm_alpha/beta, ssm_a, ssm_dt.bias, ssm_norm,
ssm_out) and hybrid_layers (the interleaved schedule). attention_layer_params
gained two terms the real file has and dense accounting did not model: the
gated q projection (2*n_h*d_k) and the q/k norm vectors.

Falsified against a real model, not against itself: fed the 0.8B configuration,
the equation now reproduces the 320-tensor inventory of
Qwen3.5-0.8B-Q4_K_M.gguf EXACTLY — 752,393,024, both layer kinds matching
tensor for tensor.

QE2E-INV-001 is still NOT asserted, and no range was widened to make it pass.
The 9b descriptor now yields 8,344,907,136, up from 8,208,519,168 but still
0.655B below [9.0B, 9.2B]. The remaining gap looks like descriptor drift rather
than missing arithmetic: 9b keeps inner_size 2048 — the value the 0.8B uses at
hidden_dim 1024 — while quadrupling hidden_dim, and its group_count 8 fails
8 * 128 == 2048, a consistency the measured 0.8B satisfies at 16 * 128. Only a
real Qwen3.5-9B file can settle it; none is on this box.

Refs #3346, #3347

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…osed

The note said the obligation was undischarged because ModelConstraints does not carry the DeltaNet shape keys. It does now, and the 0.8B count reproduces the real GGUF exactly. What actually blocks the obligation is descriptor drift at the 9b variant.

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd pv extract shrank the graph by 356 triples instead of refusing

d3cc76f rewrote the note on the QE2E-INV-001 binding and dropped the
trailing `"` — contracts/binding.yaml stopped being valid YAML at line 764
(`found unexpected end of stream`). Nothing in the PR noticed because
`pv extract contracts` does not refuse a binding registry that will not
parse: it emitted a graph with 15,244 triples where main has 15,600 —
every bound symbol AFTER the broken entry (prune::run, distill::run,
harness_ir::*, ptx_explain::run, …) silently gone — and `--check`
would have agreed with itself. Found while regenerating the derivative
for this adoption, by the drop, not by any gate.

One character. With it, binding.yaml parses (156 entries, same as main)
and the extraction is byte-identical to main's committed contracts.nt, so
this PR owes no graph change after all.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift
noahgift force-pushed the PMAT-3346-gdn-shape-keys branch from d3cc76f to da1ee15 Compare September 20, 2026 22:41
@noahgift

Copy link
Copy Markdown
Contributor Author

Adopted and rebased — and the rebase found a real defect in the branch, now fixed.

What was done. git rebase --onto origin/main c5b7bce70 with the by-id roadmap driver bound: the first branch commit was byte-identical (git patch-id) to main's squash of #3348, so only this PR's own three commits were replayed. No conflicts. New head da1ee1506.

What the rebase found. Regenerating contracts/contracts.nt on the rebased head (the R-18 derivative rule) produced a graph 356 triples smaller than main — and the missing nodes were prune::run, distill::run, harness_ir::*, symbols this PR never touches. Cause: d3cc76f1f rewrote the QE2E-INV-001 note in contracts/binding.yaml and dropped the closing ", so the file stopped being valid YAML at line 764 (found unexpected end of stream). Every bound symbol after that entry vanished from the extraction. pv extract did not refuse; it emitted the smaller graph, and --check would have agreed with itself. One character, fixed in da1ee1506; binding.yaml parses again (156 entries, same as main) and the extraction is now byte-identical to main's committed graph, so this PR owes no contracts.nt change after all.

Verified on the new head: format::model_family + format::model_arithmetic 114 passed; apr-cli oracle 218 passed; ont4b_shapes_gate 11 passed; roadmap additive deleted=0 and fragment guard PASS; fmt and clippy -p aprender-core -D warnings clean; census unchanged. Measured with a pv built from this tree under a pinned target dir, and main's own graph re-derived with the same binary as the control.

Not armed: per the board ruling every PR needs a quorum receipt first; that runs next. Refs #3346.

…t to judge against

Acceptance transcribed from issue #3346 as this PR answers it (the type that
reads the descriptor was wrong, not the range or the descriptor), with the 9B
range instantiation explicitly out of scope until a real 9B GGUF exists.

Refs #3346

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift

Copy link
Copy Markdown
Contributor Author

quorum-review (AD-04): NOT agreed (auto_merge: checked=true was_armed=false disarmed=false)

{
 "ticket": "PMAT-3346",
 "head": "1ffdccf2500116a4dd39af6a66a697f7bb841136",
 "width": 3,
 "executor": "agy",
 "agreed": false,
 "auto_merge": {
  "checked": true,
  "was_armed": false,
  "disarmed": false,
  "note": "auto-merge not armed"
 },
 "lanes": [
  {
   "lane": 1,
   "verdict": "FAIL",
   "findings": 2
  },
  {
   "lane": 2,
   "verdict": "NO-VERDICT",
   "findings": 0
  },
  {
   "lane": 3,
   "verdict": "PASS",
   "findings": 4
  }
 ]
}

…iases

Found by the AD-04 quorum on #3350 (lane 1, gemini-3.1-pro-high, cited
model_arithmetic.rs:144): the projection term used q_out for a gated
family's q matrix (2*n_h*d_k — attn_q emits the output gate, MEASURED in
Qwen3.5-0.8B) while the bias term still used q_dim. A bias narrower than its
projection is not a model. One token: q_dim -> q_out in the bias sum.

Why a delta-0 measurement did not catch it: no shipped family exercises the
case. Qwen3.5 has no attention bias; Qwen2.5 has biases but is not gated, so
q_out == q_dim there. The test that pinned 12 (a q_dim bias under a q_out
matrix) now asserts 16 and says why, and a second test holds the other
polarity — a non-gated family with biases is unchanged at 12.

115 model_arithmetic + model_family tests pass; oracle 218; clippy clean.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift

Copy link
Copy Markdown
Contributor Author

Quorum round 0: 1 FAIL / 1 no-verdict / 1 PASS — and the FAIL was a real finding, now fixed in e060c7e2c. Lane 1 (gemini-3.1-pro-high, cited model_arithmetic.rs:144): the projection term used q_out for a gated family's q matrix (2·n_h·d_k, measured in Qwen3.5-0.8B) while the bias term still used q_dim — a bias vector narrower than the projection it biases. One token (q_dim → q_out in the bias sum). It changed no shipped model's count — Qwen3.5 has no attention bias, Qwen2.5 has biases but is not gated — which is exactly how it survived a delta-0 measurement. The test that pinned 12 now asserts 16 and says why; a second test holds the other polarity (a non-gated family with biases stays at 12). 115 tests pass, oracle 218, clippy clean. Re-running the quorum on the new head. Refs #3346.

@noahgift

Copy link
Copy Markdown
Contributor Author

quorum-review (AD-04): three PASS — agreed (auto_merge: checked=true was_armed=false disarmed=false)

{
 "ticket": "PMAT-3346",
 "head": "e060c7e2c22bc221b5e454a9b677d09d2b27c2e2",
 "width": 3,
 "executor": "agy",
 "agreed": true,
 "auto_merge": {
  "checked": true,
  "was_armed": false,
  "disarmed": false,
  "note": "auto-merge not armed"
 },
 "lanes": [
  {
   "lane": 1,
   "verdict": "PASS",
   "findings": 5
  },
  {
   "lane": 2,
   "verdict": "PASS",
   "findings": 0
  },
  {
   "lane": 3,
   "verdict": "PASS",
   "findings": 2
  }
 ]
}

Round 0 was 1 FAIL / 1 no-verdict / 1 PASS and the FAIL was real (the bias
width, fixed in e060c7e). Round 1 on the fixed head: 3/3 PASS,
gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@noahgift
noahgift enabled auto-merge September 20, 2026 23:40
@noahgift

Copy link
Copy Markdown
Contributor Author

Cross-inspection of docs/audits/quorum-PMAT-3346.json lanes (non-author: aprender-04, 23:46Z)

lane conversation status verdict findings refs to sibling lanes / receipts / $WORK duration
lane 1 f3ea1f1c SUCCESS PASS (structured_output) 5 (2 cited) 0 644 s
lane 2 3f01ebe3 SUCCESS PASS (structured_output) 0 (0 cited) 0 55 s
lane 3 81760200 SUCCESS PASS (structured_output) 2 (0 cited) 0 394 s

Distinct agy conversation ids: 3/3; no lane references a sibling lane, another receipt or $WORK. Read from /tmp/claude-1000/-home-noah-src-aprender/51506a7a-265a-4e5c-bc53-50274de4a757/scratchpad/lanes-3346-round1 on this box. Note: round 0's pro lane found the gated-family bias-term defect (model_arithmetic.rs:144), fixed in e060c7e; this is round 1 on the new head; d4 armed via pmat-merge on the parent rule and this inspection is confirmatory.

Verdict line: 3/3 PASS, independent. Arm confirmed.

@noahgift

Copy link
Copy Markdown
Contributor Author

Folded into the 0.69 release batch #3669 by the cop (08:05Z, operator: "most PRs can be batched"). One CI run and one queue slot for all of them, and the generated files (roadmap.yaml, census, graph, shapes, README count) regenerated once. This PR's receipt is in the batch tree unchanged, and its closing keywords are carried in #3669's body. Disarmed here so the queue doesn't take it twice. It closes as landed-in-#3669 when the batch merges. Don't push here; changes go to release/0.69-batch.

@noahgift

Copy link
Copy Markdown
Contributor Author

Landed in #3669 (squash a877fa056, merged 2026-09-21T10:50:29Z). This PR's receipt stands as the constituent review; the batch folded its commits verbatim. — cop

@noahgift noahgift closed this Sep 21, 2026
@noahgift noahgift mentioned this pull request Sep 21, 2026
@noahgift
noahgift deleted the PMAT-3346-gdn-shape-keys branch September 23, 2026 16:15
guyernest pushed a commit to guyernest/aprender that referenced this pull request Sep 29, 2026
…run, one queue slot (paiml#3669)

* fix(pv): --table panicked on the real corpus — it cut a property by BYTE index (#3338)

`pv proof-status contracts/ --table` exits 101 on `contracts/`:

    byte index 40 is not a char boundary; it is inside '∈' (bytes 39..42)
    of `for Q4_K_M Qwen2.5-Coder, quantization ∈ {Q4_K, Q6_K}`

The column width is a byte count (`property.len()`, capped at 40) and
`truncate` sliced `&s[..max]`, so any property whose byte 40 lands inside a
multi-byte char panics. Eight contracts in `contracts/` do; the first one
walked is `apr-inspect-quantization-v1.yaml`.

The budget stays a byte budget — the table is laid out in bytes — and the
cut now walks back to the nearest char boundary.

Two tests, both RED before this commit (each panicked at
obligation_matrix.rs:169): the helper row, and one through
`format_obligation_table` because that is the path the operator hit. The
helper fixture asserts `!s.is_char_boundary(40)` first, so it cannot
silently stop proving anything.

Pmat-Ticket: PMAT-3347
Refs #3347, #3338

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(pv): the L2 column read an INDEX, not a link — so 3,573 of 3,753 obligations ticked (#3347)

`obligation_matrix` computed `l2_tested` as `idx < falsification_tests.len()`.
Obligation 3 was "tested" because the contract happened to hold 4 tests,
whoever those tests were about. All 7 obligations of
`qwen35-e2e-verification-v1` showed ✓ before a single test existed. The
fallback was a substring match between the obligation's `property` prose and a
test's `rule` prose, which is an inference, not a claim.

WHICH LINK, decided by counting the corpus (1,842 files, 3,792 obligations,
4,691 falsification tests), not by preference:

  proof_obligations[].discharged_by   89  (62 `falsification_tests[N]`, 13 a
                                           test id, 1 a YAML sequence, rest
                                           comma lists / prose / a kani id)
  falsification_tests[].binds_to      38
  falsification_tests[].obligation    26  (12 name an obligation id, 6 the
                                           exact property text, 8 dangle)
  proof_obligations[].id             180 of 3,792
  kani_harnesses[].obligation      1,918 — that is the L3 column, not this one

All three L2 spellings are DECLARATIONS by the contract author, so all three
are read. `binds_to` is a serde alias of `obligation`, which is safe only
because no entry carries both keys — checked, because serde turns that into a
`duplicate field` parse error rather than a silent pick.

Not read: `applies_to`. 12 `binds_to` values match one, but `AppliesTo` is an
enum with `#[serde(other)] Other`, so the string is discarded at parse and
there is nothing left to compare. Those 12 report `?`, not a false ✗.

WHAT THE COLUMN NOW SAYS. Three values, because "no test covers this" and
"nothing here says which test covers what" are different facts:

  ✓ Tested   a test in this contract cites this obligation
  ✗ Untested the contract's links resolve and none names it — or it ships
             no falsification test at all, which is a reading, not a gap
  ? Unknown  no readable link; not measured. An unread window is Unknown,
             never a tick

MEASURED over `contracts/`, and the drop IS the point — it is what the old
column was hiding:

              before   after
  L2 ✓         3,573      86
  L2 ✗           180      65
  L2 ?             0   3,602
                       (3,753 obligation rows, 873 contracts)

Nothing was adjusted to keep the number up, and no threshold was added.

The two link fields did not exist on the structs, so both keys were written to
disk and silently dropped on parse — the shape of #3314 (`id`) and #2465
(`test_harness`). `discharged_by` is typed `Citation` (scalar | comma list |
YAML sequence) because `Option<String>` failed the WHOLE corpus on
`publish-manifest-v1`: `invalid type: sequence, expected a string`.

RED first: `l2_does_not_tick_for_an_obligation_no_test_cites` — two
obligations, two tests, both citing OB-A. It asserted through the rendered
table (the surface that was lying, and API-stable), so it is the SAME test
before and after: on the old code `Beta holds | ✓`. Its OB-A arm keeps a fix
that merely stopped ticking everything from passing.

Out of scope_paths, and forced: `lint/strict_test_binding.rs` holds the only
EXHAUSTIVE `FalsificationTest` literal in the tree (no `..Default::default()`),
so no schema field can compile without that one line.

Pmat-Ticket: PMAT-3347
Refs #3347, #3091, #3114

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(pv): single-file --strict-test-binding reported every ref missing — it now refuses (#3347)

The gate resolves cited test names against a source index rooted at the
contract path's PARENT. The directory form gets the repo root and finds
`crates/`; the single-file form gets `contracts/`, which holds no source, so
every cited ref resolves to nothing and all of them are reported missing.

Measured on `contracts/pv-artifact-kinds-v1.yaml`, a control whose eight refs
all resolve:

    pv lint contracts/pv-artifact-kinds-v1.yaml --strict-test-binding
        total_refs 8, existing 0, missing 8
    pv lint contracts/ --strict-test-binding
        total_refs 548, existing 521, missing 27  <- none of the 27 is this one

A control contract failing identically to a broken one is a gate that cannot
discriminate, and it fails SILENTLY: `passed` is true in non-strict mode, so
the run still says `Result: PASS` while printing eight false findings.

REFUSED rather than repaired. The scan root is computed in
`provable_contracts::lint::run_lint`, outside this ticket's scope; a refusal
lives in the caller, is honest, and cannot be mistaken for a clean bill. If
the root is later made explicit (a `LintConfig` field the CLI can feed, which
is what `--crate-dir` does NOT do today), this refusal is what should be
deleted.

Exit 1, not the exit-2 `decline:` class: exit 2 belongs to `ZeroContracts`,
whose message ("0 contracts under ...") would be false here — there IS a
contract; it is the gate that cannot run over it.

Two tests, in the CI-wired `cli_integration` target: the refusal names the
flag and prints no PASS and no findings, and a directory-form control proves
the refusal is specific to the single-FILE form rather than to the flag.

Pmat-Ticket: PMAT-3347
Refs #3347

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(crux): AutoGluon becomes a CRUX competitor — category O, 24 contracts, 25 tickets, registry edit + mutation proof

Research of ../autogluon (1.6.3 @ 77946149) as a competitive-research
source for aprender, landed the way category N landed linfa and burn
(#3169): the competitor is admitted to the CLOSED registry, every story
is a real contract with falsification gates, every story is a registry
row, and every missing/partial story has a GitHub issue and a roadmap
fragment.

What AutoGluon is, measured from the tree rather than recalled: three
predictors. TabularPredictor (65 public methods, 11 presets, 24 model
families of which 8 are tabular foundation models added in 1.4-1.6),
TimeSeriesPredictor (Chronos-2/Toto-2 pretrained, 30+ local/deep models,
16 metrics incl. WQL/MASE/RMSSE, auto backtesting since 1.5) and
MultiModalPredictor. Evidence under evidence/crux/autogluon/.

What aprender has, measured at eb262f8eb: automl/ is a single-estimator
hyperparameter tuner (TPE, grid, random, DE, TimeBudget, EarlyStopping);
time_series/ is one univariate ARIMA; encoders, calibration, SHAP/LIME/
permutation importance and KFold/cross_validate exist as building
blocks. No predictor-level fit(label), no leaderboard, no bagging,
stacking or greedy weighted-ensemble selection, no panel forecasting,
no quantile forecast metrics. The gap is the AutoML UX, not the
algorithms.

Category O — AutoML Parity — 24 stories: 9 P0 (the README hello-world:
fit(label), problem-type inference, presets, leaderboard, feature
pipeline, weighted ensemble, time budget, panel forecaster, quantile
metrics), 9 P1 (bagging, stacking, importance, threshold calibration,
deployment artifact, tabular foundation model, backtesting, local
baselines, pretrained forecaster), 6 P2 (refit_full, distill,
infer_limit, fit diagnostics, memory-aware fit, covariates).
MultiModalPredictor, autogluon.cloud, MLZero and Ray-parallel fits are
CUT on the epic with reasons.

  CRUX_COMPETITORS: [&str; 14] -> [&str; 15]   + autogluon

Not a BEAT pillar: aprender claims no pinned-benchmark win over
AutoGluon. Both registry tests that keep BEAT_INCUMBENTS and
CRUX_COMPETITORS apart are extended, not worked around.

Tickets: epic #3370, stories #3371-#3394, label pareto-autogluon.
Roadmap: 25 fragments under docs/roadmaps/entries/, roadmap.yaml
regenerated by the aggregator (idempotent check passes).

Spec: docs/specifications/crux-competitive-research-ux-workflows.md
v2.2 -> v2.3 — §3 gains rows for linfa+burn (category N, which #3169
never recorded there) and AutoGluon; §5 gains Category O; §6 notes
that coverage_intake in the YAML is the source of truth.
coverage_intake 267 -> 291 (partial 72 -> 77, missing 152 -> 171).

Verification:
- pv built from THIS tree validates 25/25 (24 new + master). The stale
  ~/.cargo/bin/pv rejects crux-O-01 with CRUX-002 — the behavioural
  delta proves the registry edit engaged.
- Mutation-verified: deleting "autogluon" from CRUX_COMPETITORS turns
  competitor_registry_covers_the_corpus_vocabulary RED with "autogluon
  is used by contracts/ and must stay in CRUX_COMPETITORS" and
  the_real_crux_registry_rows_are_all_in_domain RED. Restored: 20/20.
- cargo test -p aprender-contracts --lib: 1526 passed, 0 failed.
- Every falsification gate is LIVE-PENDING prose (no `::`), so
  strict-test-binding has nothing to refuse; the obligations are
  RECORDED as unfalsifiable-by-absence, not satisfied.
- README CONTRACT_COUNT regenerated 1835 -> 1866 by readme_sync.sh.
- Guards: roadmap fragment/ids/sorted/additive/completion, contract
  test-binding and enforcement, shell-lint ratchet, hardcoded paths,
  readme claims, grep -q ratchet — all rc=0.

Pmat-Ticket: PMAT-3370

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(crux): bind the admission PR to its own ticket PMAT-3401 — the epic's acceptance criteria describe the programme, not this diff

Round 2 of the quorum read the epic (PMAT-3370) as the ticket and refused
the admission for not implementing the 24 stories. The admission is its
own unit of work with its own done-when; this fragment says so.

Closes #3401
Pmat-Ticket: PMAT-3401

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(crux): PMAT-3401 title inventories the diff exactly — 26 fragments (its own included) and the scaffold CATEGORY_NAMES edit

Quorum round 3 lane 1 refused on two literal mismatches between the
ticket and the diff: '25 roadmap fragments' (there are 26 once this
ticket's own fragment lands) and an unlisted edit to
scripts/crux_scaffold_contracts.py. The title now lists every path.

Pmat-Ticket: PMAT-3401

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* roadmap(PMAT-3495): VERIFY-001 on 0.71.0 — run the 132 Kani harnesses in CI (proof credit from runs, not declarations), then a Verus pilot on one dequant/parser function (#3495)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* quorum(PMAT-3496 brief, PR #3496): 3/3 PASS on 57e0e4dfb — gemini lanes, measured; author claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* PMAT-3577: the logit-parity receipts under contract — parity-receipt-v2, extract:parity-receipt, and the 7 back-filled

Until this row the logit-parity records under evidence/parity/** had NO validator
of any kind. Not a weak one — none. That is why seven of them sat in the tree
carrying no comparator for months: there was nothing that could have noticed.

WHAT LANDS
  · contracts/parity-receipt-v2.yaml — three shapes: parity-receipt-complete
    (closed, ignoredProperties empty), parity-comparator-self, -oracle. The
    subset has no sh:or, so the comparator split is two shapes over two
    subclasses the extractor assigns by kind.
  · contracts/parity-receipt-v1.yaml — the retired layout, recorded with NO
    shape: all instances were migrated, and a shape whose target class nothing
    instantiates passes vacuously. The enforcement that replaces it fires — the
    extractor refuses an unmigrated record BY NAME, and so does the predicate.
  · ontology/extract/parity_receipt.rs — record / unmigrated / other, and
    skipping is never silent.
  · scripts/parity_receipt_denominator.sh + evidence/parity/EXPECTED_RECEIPTS —
    the count is pinned by an INDEPENDENT predicate. An extractor checked
    against a number the extractor produced proves nothing.
  · Unknown{ExtractorMiss}, exit 2 — a new element of the verdict lattice. An
    extractor that silently saw the wrong corpus reports the same "no
    violations" as one that saw all of it.
  · the 7 records migrated to v2 and back-filled with comparator {kind: self,
    reason}, in this commit, as the row requires.

THE DENOMINATOR IS 7, NOT 8. The #3574 receipt is in PR #3575, still open; it is
not on main. Measured: 113 files under evidence/parity, 7 parity records, 0 with
a comparator. #3575 bumps it to 8 when it lands — this row's own falsifier on its
first real use.

check_parity_receipt.sh IS NOT TOUCHED, and that is the finding, not an omission.
It validates the THROUGHPUT family (instrument, protocol_ref, lanes[],
decode_tok_per_sec, the #2696 cross-class defect); a logit record has never
carried one of those keys. Folding it in — as item 6 asked — would have deleted
the #2696 validator from a family nobody was watching. One validator per artifact
family, and the discriminator is the artifact's required keys, never its filename.

THE BACK-FILL IS A RELABEL. `raw` is the original apr parity --json document key
for key; every envelope field is quoted from committed evidence named in each
record's provenance.record. model_sha256 is carried only by the one record that
measured it: hashing the files today and attaching that to a receipt about
2026-09-06 would be a claim about a different world wearing a witness's clothes.

One derived field was wrong first time and is worth recording: result.verdict
copied raw.parity — apr's own per-position flag — which says PASS for the two
1.5B cells their own RECORD.md calls RED. It now resolves the threshold from
thresholds.yaml and reproduces all seven readings the RECORD.md files state,
both REDs included. No threshold is ever typed into a shape.

NOTHING IS ARMED. armed_shapes lives in lint-baseline.json, a shared file this
row may not touch (decision 7); the shapes are computed and reported, as
ladder-green was at ONT-4c1. Arming is a follow-up with the label.

CONTROLS, both directions: 7 violations with the comparator stripped from all
seven → 0 as committed; exactly 1 for a single plant, naming focus node and
property; widening sh:in to accept `oracle` turns ont4c3_parity_receipts RED, and
mutating only the real contract turns the fixture-drift test RED; an unmigrated
record declines at exit 2 naming the file; 2 receipts against a denominator of 1
declines naming both numbers; pass / fail / decline are 0 / 1 / 2.

  cargo test -p aprender-contracts --lib ontology::extract::parity_receipt  10 ok
  cargo test -p aprender-contracts-cli --test ont4c3_parity_receipts         8 ok
  bash scripts/parity_receipt_denominator.sh --self-test                     4 ok
  pv lint contracts --gate {sigma,relations,shapes}          Pass, 0 violations
  pv extract contracts --check                                            rc 0

ONT-4c3 is NOT bound in the ONT-001 ledger: that ledger is paiml/infra's
paiml-ontology.md, where v4.8 still defines ONT-4c3 as KERNEL receipts. The
re-scope is an infra PR, not this one — raised with the cop rather than left as a
checked box. Receipt: docs/audits/impl-PMAT-3577-receipt.md.

Refs #3577, #3576, #3269, #3575, #3567

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: name it WrongCorpus, not ExtractorMiss — a prefix of its opposite is a defect

`ExtractorMissing` already meant the opposite thing: an extractor that does not
exist. `ExtractorMiss` would have sat beside it in the same 18-element lattice,
one letter apart, with the shorter a PREFIX of the longer — `grep ExtractorMiss`
matches both, and any substring test over the reasons merges them silently. That
is the defect class this repo keeps paying for (#3573 today), so the name now says
what happened: the extractor RAN and read a corpus the tree does not declare.

Renamed before it reached main, when a rename is a sed and not a migration.
Caught in review by the cop.

Refs #3577, #3573

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: the ONT-4c3 divergence is filed as paiml/infra#814, not left in a thread

Two repos holding different definitions of one identifier is a row, not a flag in
a message: nothing collides where a tool would see it, so it collides in a
person's head months later when they implement the ledger's meaning and find
their correct work unusable.

Refs #3577, paiml/infra#814

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* #3605: removed_by gets a shape — refusal-receipt-v1, with a closed sentinel set so the required field manufactures nothing

`removed_by` had 0 occurrences tree-wide, no schema and no validator. Every
refusal #3597 writes would have minted an unenforced convention, and a tree full
of consistent-looking `removed_by:` lines reads as validated when it is
decoration.

PHASE 0 — ITS OWN CONTRACT, AND NOT BECAUSE IT IS EASIER TO WRITE

Option B is refused on a fact: `parity-receipt-v1` DOES NOT EXIST AT HEAD. It is
in unmerged #3600, and there it is deliberately the RETIRED layout with no shape
and no instances. A live validated field on a superseded contract with zero focus
nodes is a field nothing can carry. Independently, by the discriminator this tree
has now used twice — an artifact family is named by its REQUIRED KEYS, never by
its filename — a parity receipt requires host, backend, comparator,
threshold_source and per-position metrics; a refusal requires a verb, a reason,
an exit code and removed_by. They share no required key.

A THIRD option was considered and refused, and it is the more attractive one:
put removed_by on apr-cli-commands-v1.yaml, where the verb already IS the focus
node and the universe is already the trustworthy 111. Refused because the
registry is the UNIVERSE and #3597's method is to DIFF the registry against the
buckets. If the buckets live in the registry, the denominator and the numerator
are the same artifact and the diff is vacuous by construction. The registry says
what EXISTS; a refusal says what was TRIED. Keeping them apart is what lets that
diff be a real diff.

THE FORCED-BINDING TRAP, AND WHY minCount 1 IS SAFE HERE

A required field with no escape manufactures false data: `pv validate` requires
kani_harnesses so authors fabricate one, and apex#57 had lean_theorem copied
verbatim into twelve contracts resolving to nothing, gate green throughout. The
escape is a closed sentinel set, so "there is legitimately nothing here" is
SAYABLE and still CHECKABLE:

  v<major>.<minor>   the release that removes it
  never              refused permanently by design
  unscheduled        a defect with no release chosen

tbd, soon, n/a, pending, a bare `0.70`, a patch-level `v0.70.1` and a git sha are
all RED. A version rather than a sha because a refusal answers "which release do
I need?" and a sha is precise about the tree and silent about the boundary; the
sha form is made INVALID rather than discouraged, because a shape that permits
two spellings gets both.

BOTH ARMS, and the second is the point: 9 cases over 8 fixtures. refusal-ok
carries all three declared forms and passes with 3 focus nodes;
refusal-undeclared-sentinel plants the PLAUSIBLE `tbd` and must go red. Mutation
proved red-capable: widening the pattern to accept tbd fails exactly
a_plausible_but_undeclared_sentinel_is_refused and
accept_and_refuse_are_distinct_answers, and restoring makes 9/9 green. Fixtures
carry the real contract byte for byte, so widening it without them is caught too.

No new extractor: entity {type: json, ref} + vocabulary is the existing machinery
for "validate this document", and a second reader for one more family is the
thing this tree keeps filing against.

done_when 6 — the interim re-check list is EMPTY: `removed_by` still has 0
occurrences at HEAD, so no refusal written under the interim needs revisiting.

The ledger is seeded with the one refusal already measured (`apr bench` refuses
qwen35) so the shape has a real focus node instead of passing vacuously over an
empty list. WHICH verbs land there is #3597's bucket, not this row's.

Refs #3605, #3597, #3600, #3080

Pmat-Ticket: PMAT-3605

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: count parity-receipt in by_entity_type — ONT-001 v4.10's probe reads ABSENT otherwise

The spec's probe asks `by_entity_type["parity-receipt"] == 7`. Measured on this
branch before the change: the map carried pv-contract, gguf, apr-model, code and
lean, and NO parity-receipt key. The seven records were there — by_shape showed
parity-receipt-complete=7 — but the entity-type map did not carry them, so the
probe would have read ABSENT.

AN ABSENT KEY IS NOT ZERO. A consumer treating it as one measures nothing and
calls it a pass — the same shape as #3610, one map over.

Registering the entity type in Sigma was not enough; it has to be counted where
the probe looks. A test now asserts the key exists AND that it equals the shape's
own focus-node count, so the two numbers cannot drift apart.

Refs #3577, #3610, paiml/infra#831

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: regenerate contracts.nt with a pv built from THIS worktree — the shared target dir served another branch's binary

The merge commit's contracts.nt was 107 triples short: every parity-receipt and
parity-comparator node was missing, because `cargo metadata`'s target_directory
is shared across worktrees and the pv on PATH had been built from a different
branch minutes earlier.

AND `pv extract contracts --check` PASSED ON IT, because the check re-derives the
graph with the same binary. A stale tool comparing an artifact against its own
re-derivation agrees with itself about nothing being there — the derived file and
the checker were wrong in the same direction, which is the only way that gate can
fail to fire.

Rebuilt with CARGO_TARGET_DIR pinned to this worktree; the 107 triples return and
by_entity_type[parity-receipt] reads 7 rather than absent.

Refs #3577

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* register the new CLI test in scripts/tree_reader_tests.txt

The registry is derived from the tree and drifts the moment a test that reads the
tree is added without listing it. Each of the three new tests failed this guard on
its OWN branch — not a shared commit, as first read.

Pmat-Ticket: PMAT-3598

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* register the new CLI test in scripts/tree_reader_tests.txt

The registry is derived from the tree and drifts the moment a test that reads the
tree is added without listing it. Each of the three new tests failed this guard on
its OWN branch — not a shared commit, as first read.

Pmat-Ticket: PMAT-3598

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* #3600: the migration broke TWO readers of the records, including the release judge

I asked "who else reads this?" for the <unk> assumption and did not ask it for my
own migration. Two consumers read the legacy top-level metrics:

  crates/apr-cli/src/commands/parity_admission.rs  -> mac-check RED
  scripts/check_model_parity.sh                    -> C14, the RELEASE judge

The second is the one that matters: C14 is the pre-publish dogfood's parity gate,
and on the migrated records it reported "no per-position metrics in the output"
for a record that is fine. The migration would have taken the release gate down.

THE RULE, STATED ONCE IN BOTH READERS: a v2 receipt EMBEDS the raw
`apr parity --json` document under `raw`, so look inside the envelope when there
is one. A fresh `apr parity` run is the raw document itself and carries the
readings at the top level. Those are two different INPUTS — a tool's output and an
archived receipt quoting it — not two spellings of one, which is the distinction
that makes this a rule rather than the permissiveness #3613 refuses.

The self-test's fixture builder read the same way, so the must-RED twin was being
built from a KeyError and that control could not have fired.

Verified: 7B sentinel PASS, 1.5B sentinel RED (unchanged from pre-migration),
check_model_parity.sh --self-test 26/26, parity_admission 19 tests.

Refs #3600, #3577

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3604: the F2 hybrid guard runs once per (model sha256, apr version, device) and leaves a receipt

f2_validate_qwen35 proves the CUDA hybrid forward against a 64-position CPU
reference before it will serve a token, on EVERY apr run. Measured (#3598
row 1): 67 % of a 14 s time-to-first-token, 90-93 % of it the CPU forward.
The guard is right to exist and wrong to run per call.

Now: the first run of a (model sha256, apr version, device) triple validates
and writes a receipt; a later run whose triple matches reads it and skips the
forward; `apr run --revalidate` forces a fresh run and rewrites it.

MEASURED HERE, RTX 4090, Qwen3.5-4B-Q4_K_M, the 144-word row, --max-tokens 1,
GPU occupancy recorded before every run, binary built from this tree under a
PINNED target dir (the shared one handed me another worktree's binary first):

  --revalidate, warm cache   17.18 s wall   guard 9,747 ms on 65 positions
  receipt hit, warm cache     7.39 s wall   guard 0 ms      sha256 1,173 ms

  -9.8 s wall; guard 9,747 -> 0; the key costs 1.17 s/run (2.74 GB at
  2.3 GB/s) and is printed separately so it cannot hide in either number.

THE RECEIPT IS THE VALIDATION, WHICH IS WHY IT IS STRICT. Every path that is
not "three keys match" validates, and the three planted-receipt falsifiers
were run END TO END in the real binary, not only as unit tests:

  wrong model sha256   -> re-validated: "receipt is for model bbbb…, this file is 00fe…"
  wrong apr version    -> re-validated: "receipt written by apr 0.61.0, this is 0.68.2"
  wrong device         -> re-validated: "receipt written for NVIDIA GB10, this device is …4090"
  corrupt file         -> re-validated: "receipt unreadable (…: not a receipt)"
  missing file         -> re-validated: "no receipt for this model"

Absence is never consent, and absence and unreadability are told apart.

A RECEIPT IS WRITTEN ONLY AFTER A VALIDATION THAT JUDGED SOMETHING. The guard
has three early exits that let the GPU serve without comparing a position —
SKIP_PARITY_GATE=1, a probe under two tokens, a CPU reference that would not
run. f2_validate_qwen35 now returns F2Verdict {Accepted{positions_judged},
Rejected, NotJudged} instead of bool, and only Accepted writes; otherwise a
one-token prompt would "validate" the triple for every prompt after it.

The decision table (f2_receipt.rs) is pure and CUDA-free, so its 13 tests run
on every build. --revalidate reaches the guard through the same env seam the
guard already reads SKIP_PARITY_GATE from, rather than a 38th positional
parameter on run_entry::run and six forward signatures #3606 is changing.

done_when 5 is partial and says so: [source=receipt|fresh] is on the guard's
stderr line; the `apr run --json` field lands with #3606's StageTimings, and
F2Outcome{source, validate_ms, sha256_ms, receipt_path} is returned to the call
site for exactly that. #3606 and this PR both edit f2_validate_qwen35's return
path; whichever lands second reconciles ~10 lines, and #3606's lane is told.

Evidence: evidence/perf/3604/MEASUREMENT.md + the stderr of all nine runs.

Refs #3604, #3596, #3598, #3606

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* roadmap: mint PMAT-3604 so the AD-04 quorum for #3634 can run

pmat work status PMAT-3604 refused ('Item not found'): #3604 was minted as a
GitHub issue with a done_when but never as a roadmap entry, and the quorum
script hard-requires the work item. Fragment + aggregate, nothing else.

Refs #3604, #3634
ont-delta: none — a roadmap entry; no ontology entity, shape, reason or resolution

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(run): --gpu that fell back to CPU reported success — the refusal had no caller (#3602)

`apr run --gpu` on a model whose GPU attempt is rejected at runtime printed a
result and exited 0. Measured on lambda (RTX 4090 sm_89, apr 0.68.2 e6f77c98c,
qwen2.5-coder-0.5b-instruct-q4_k_m):

  exit=0
  stderr: warning: GPU output diverges from CPU at position 1 (cosine 0.4153)
  stdout: { "used_gpu": false, "inference_time_ms": 33646.13 }

33.6 seconds on the CPU, reported as a successful --gpu run, with `used_gpu:
false` as the only signal — the same value a deliberate CPU run reports.

The decision was already made, recorded and unit-tested. `registry::after_generation`
implements it (R-0b, #3002/#3042): a FORCED accelerator that fell to CPU is a
refusal, a DEFAULT selection that fell to CPU gets a corrective line. A git grep
found its only callers were its own tests. `registry::announce` and
`registry::parity_line` are in the same state, which is why a real run prints
zero `selected:` and zero `parity:` lines.

So this is not a new policy. It is the recorded one, reached from `apr run`:

- `dispatch.rs` classifies the request ONCE via `registry::Request::wanted()`
  rather than re-deriving "forced" — two spellings of one rule drift apart.
- `run_entry::run` calls `reconcile_accelerator` before any success output.
- `--json` gains a `backend` object: `requested` / `ran` / `fell_back`, because
  `used_gpu: false` alone collapses "ran on CPU deliberately" with "asked for
  the GPU and was refused it".

`accel::tests::every_accelerator_surface_calls_the_refusal` passed throughout:
it guards `ensure_available`, the BUILD-time refusal, not the RUNTIME one. A
guard over one of two refusals reads as coverage of both.

Falsifier, both directions (`run_tests_accel_reconcile.rs`, 6 cases). Planting
the pre-fix behaviour (`reconcile_accelerator` → `Ok(None)`) turns
`a_forced_accelerator_that_ran_on_cpu_is_refused` and
`the_json_distinguishes_a_deliberate_cpu_run_from_a_rejected_gpu_run` RED while
`a_forced_accelerator_that_actually_ran_on_gpu_says_nothing` stays green — so
the tests discriminate rather than just failing. `used_gpu: None` is Unknown,
not a fallback: refusing on it would make every non-reporting path a hard error.

NOT included, deliberately: the rejection's REASON (cosine, position) is on
stderr but not in the JSON. It is produced inside realizar's F2 gate and no
channel carries it to the CLI; adding one is a #3606-shaped follow-up, not
something to approximate with a guess here.

Refs #3602, #3483
Pmat-Ticket: PMAT-3602

* test(PMAT-3346): the measured Qwen3.5-0.8B inventory the dense path cannot reproduce

QE2E-INV-001 could not be judged because nothing in the tree held a MEASURED
Qwen3.5 tensor inventory to judge against. This adds one: the 320 tensors of
~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782...dc517), read straight from
the GGUF header rather than from a model card.

It already falsifies the current arithmetic. Dense/GQA accounting applied to
that file gives 644,400,128 against a measured 752,393,024 — short by
107,992,896, 14.4% of the model, because 18 of the 24 layers are Gated DeltaNet
and no term here counts their conv, gate, state or output projections.

Two shapes in the file are also not what dense accounting predicts, and both are
pinned: attn_q is [1024, 4096] = 2 * num_heads * head_dim (the q projection
emits the attention output gate alongside the query; attn_output [2048, 1024]
confirms num_heads * head_dim = 2048), and attn_q_norm/attn_k_norm are present
at head_dim. The file is also TIED — it has no output.weight.

Refs #3346

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(PMAT-3346): ModelConstraints carries the gated-DeltaNet shape, so a hybrid model can be counted

contracts/model-families/qwen3_5.yaml declares inner_size, state_size,
conv_kernel, group_count and full_attention_interval under constraints:, and
ModelConstraints carried none of them. A Gated DeltaNet layer's parameters live
entirely in those dimensions, so every consumer of the descriptor counted
Qwen3.5 as if three quarters of its layers did not exist.

Carried through as ModelConstraints::deltanet: Option<DeltaNetShape> — the
runtime YAML loader (parsing.rs) and the compiled-in registry (build_parsing.rs
+ build_codegen.rs) both populate it, and FALSIFY-MF-QWEN35-010 pins that the
declared values survive the trip and that no other family acquires a shape it
never declared. qwen3_5.yaml is the only descriptor with these keys, so every
other family keeps byte-identical accounting.

model_arithmetic gains gated_deltanet_layer_params (one term per GGUF tensor:
attn_qkv, attn_gate, ssm_conv1d, ssm_alpha/beta, ssm_a, ssm_dt.bias, ssm_norm,
ssm_out) and hybrid_layers (the interleaved schedule). attention_layer_params
gained two terms the real file has and dense accounting did not model: the
gated q projection (2*n_h*d_k) and the q/k norm vectors.

Falsified against a real model, not against itself: fed the 0.8B configuration,
the equation now reproduces the 320-tensor inventory of
Qwen3.5-0.8B-Q4_K_M.gguf EXACTLY — 752,393,024, both layer kinds matching
tensor for tensor.

QE2E-INV-001 is still NOT asserted, and no range was widened to make it pass.
The 9b descriptor now yields 8,344,907,136, up from 8,208,519,168 but still
0.655B below [9.0B, 9.2B]. The remaining gap looks like descriptor drift rather
than missing arithmetic: 9b keeps inner_size 2048 — the value the 0.8B uses at
hidden_dim 1024 — while quadrupling hidden_dim, and its group_count 8 fails
8 * 128 == 2048, a consistency the measured 0.8B satisfies at 16 * 128. Only a
real Qwen3.5-9B file can settle it; none is on this box.

Refs #3346, #3347

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* contracts(binding): the QE2E-INV-001 note blamed a gap that is now closed

The note said the obligation was undischarged because ModelConstraints does not carry the DeltaNet shape keys. It does now, and the 0.8B count reproduces the real GGUF exactly. What actually blocks the obligation is descriptor drift at the 9b variant.

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3346 (adoption): the QE2E-INV-001 note lost its closing quote, and pv extract shrank the graph by 356 triples instead of refusing

d3cc76f1f rewrote the note on the QE2E-INV-001 binding and dropped the
trailing `"` — contracts/binding.yaml stopped being valid YAML at line 764
(`found unexpected end of stream`). Nothing in the PR noticed because
`pv extract contracts` does not refuse a binding registry that will not
parse: it emitted a graph with 15,244 triples where main has 15,600 —
every bound symbol AFTER the broken entry (prune::run, distill::run,
harness_ir::*, ptx_explain::run, …) silently gone — and `--check`
would have agreed with itself. Found while regenerating the derivative
for this adoption, by the drop, not by any gate.

One character. With it, binding.yaml parses (156 entries, same as main)
and the extraction is byte-identical to main's committed contracts.nt, so
this PR owes no graph change after all.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(roadmap): mint PMAT-3602 as a work item so the AD-04 quorum can resolve it

`quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work`
reads the aggregated roadmap rather than an id counter. The branch carried the
trailer `Pmat-Ticket: PMAT-3602` while no such work item existed — a GitHub
issue number is not a work item.

The id is DERIVED from the issue (`pmat work add --github-issue` calls that
"the path to prefer": GitHub allocates centrally, so two agents cannot be handed
the same id, which `max(id)+1` cannot promise). Fragment, not a direct edit to
roadmap.yaml: additive, one file, and it does not make every stacked branch
dirty on the same lines.

Acceptance is hand-entered from the issue's done_when — `pmat work add` derives
neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414).

Guards: diff_additive, fragment_required, ids_unique, sorted,
completion_is_cited all exit 0; `make roadmap-aggregate-check` reports
`roadmap.yaml == aggregate(48 fragment(s)), idempotent`.

Refs #3602
Pmat-Ticket: PMAT-3602

* roadmap: PMAT-3637 — the registration row this PR is, so its quorum judges fidelity not implementation

Round 0 ran with --ticket PMAT-3604 and all three lanes FAILed for the same
reason: the diff implements none of PMAT-3604's criteria. Correct — it was
never meant to; it registers the ticket so #3634's quorum can start. The
receipt is kept as quorum-PMAT-3637-r0-misframed-as-3604.json (a FAIL is a
record, not a mistake to erase). PMAT-3637 is the row this diff satisfies:
fidelity of the PMAT-3604 entry to issue #3604's done_when, additive
aggregate, nothing outside docs/roadmaps/. Round 1 runs against it.

Refs #3604, #3634
ont-delta: none — roadmap entries only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(roadmap): mint PMAT-3605 as a work item so the AD-04 quorum can resolve it

`quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work`
reads the aggregated roadmap rather than an id counter. This branch carried the
trailer while no such work item existed — a GitHub issue number is not a work
item until something mints it.

The id is DERIVED from the issue: `pmat work add --github-issue` calls that
"the path to prefer", because GitHub allocates centrally and `max(id)+1` cannot
promise two agents different ids. A fragment rather than a direct roadmap.yaml
edit — additive, one file, and it does not make every stacked branch dirty on
the same lines.

Acceptance hand-entered from the issue's done_when; `pmat work add` derives
neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414).

Guards: diff_additive, fragment_required, ids_unique, sorted,
completion_is_cited all exit 0; `make roadmap-aggregate-check` idempotent.

Refs #3605
Pmat-Ticket: PMAT-3605

* PMAT-3346: the roadmap row, so the AD-04 quorum for #3350 has a ticket to judge against

Acceptance transcribed from issue #3346 as this PR answers it (the type that
reads the descriptor was wrong, not the range or the descriptor), with the 9B
range instantiation explicitly out of scope until a real 9B GGUF exists.

Refs #3346

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* roadmap: round 1 said two true things — fix both

Lanes 1+2 (gemini-3.1-pro-high, gemini-3.8-flash-high) FAILed PMAT-3637 on
(a) the round-0 receipt committed under docs/audits/, which criterion (3)
forbade — dropped; the receipt judged the wrong question and its ticket
field would read as a verdict on PMAT-3604; and (b) a criterion in the
PMAT-3604 notes that issue #3604 does not state ('Only an Accepted verdict
writes a receipt') — it is PR #3634's design decision, not a done_when
item; removed from the transcription. Criterion (3) now admits this row's
own receipt at docs/audits/quorum-PMAT-3637*.json, which it must, or the
PASS receipt could never be committed.

Refs #3604, #3634
ont-delta: none — roadmap entries only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(run): quorum round 1 found the --json surface skipped and a branch that could never run

Two findings from lane 1 (gemini-3.1-pro-high) of the AD-04 quorum, both
correct, both fixed. Lane 3 passed the PR without seeing either.

1. `reconcile_accelerator(...)?` ran BEFORE `print_run_output`, so a rejected
   `--gpu` run early-returned and `--json` emitted nothing at all — the exact
   surface #3602 item 1 names. Machine surfaces (`--json`, `--stream`) now emit
   before the refusal propagates; the human surface still prints nothing.

   This is a DELIBERATE deviation from `after_generation`'s contract, which says
   the caller "must print NO output" on a forced refusal. Named rather than
   quiet: that rule exists so a CPU result is never read as a GPU success, and a
   document carrying `"backend": {"fell_back": true}` beside exit 14 cannot be
   read that way, while a human-formatted success blob can. The protective half
   is kept; the half that blinded `--json` consumers is not.

2. The caller's `if let Some(note)` arm was UNREACHABLE. `announced` was
   `Some("gpu")` exactly when `forced` was true, and `after_generation`'s
   corrective-line branch requires `forced == false` — so the Some arm could
   never execute. A branch with no reachable caller, which is precisely the
   defect this PR fixes, reproduced one layer down while fixing it.

   `reconcile_accelerator` now returns `Result<()>`, wires the FORCED half only,
   and says in its doc comment why the default-selection half is not wired:
   nothing calls `registry::announce`, so there is no recorded announcement for
   a default selection, and manufacturing one would assert a choice this process
   never made. That is REG-8, not this PR.

Verification: cargo test -p apr-cli --lib = 7289 passed, 0 failed, 12 ignored;
cargo fmt --all --check = 0; cargo clippy -p apr-cli --lib -D warnings = 0.

Refs #3602
Pmat-Ticket: PMAT-3602

* PMAT-3346 (adoption): a bias vector is as wide as the projection it biases

Found by the AD-04 quorum on #3350 (lane 1, gemini-3.1-pro-high, cited
model_arithmetic.rs:144): the projection term used q_out for a gated
family's q matrix (2*n_h*d_k — attn_q emits the output gate, MEASURED in
Qwen3.5-0.8B) while the bias term still used q_dim. A bias narrower than its
projection is not a model. One token: q_dim -> q_out in the bias sum.

Why a delta-0 measurement did not catch it: no shipped family exercises the
case. Qwen3.5 has no attention bias; Qwen2.5 has biases but is not gated, so
q_out == q_dim there. The test that pinned 12 (a q_dim bias under a q_out
matrix) now asserts 16 and says why, and a second test holds the other
polarity — a non-gated family with biases is unchanged at 12.

115 model_arithmetic + model_family tests pass; oracle 218; clippy clean.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(run): quorum round 2 — a --json --benchmark leak, a comment claiming coverage it lacks, and a stray artifact I swept in

Three findings, two lanes, all correct.

1. LANE 1 — `--json --benchmark` leaked a human success blob on a refused run.
   The refusal path carried its OWN copy of "is this a machine surface"
   (`stream || output_format == "json"`) and dropped `!benchmark`. So that
   combination entered the branch, matched neither machine arm inside
   `print_run_output`, and fell through to the human benchmark rendering — for a
   run being refused. Two spellings of one condition drifting apart in the gap
   between them, which is the defect this PR's own dispatch comment warns about.

   One spelling now: `emits_machine_output(stream, output_format, benchmark)`.
   `the_machine_output_predicate_matches_print_run_output` pins it to the arms
   it describes over every flag combination; planting the pre-fix spelling turns
   it and `json_plus_benchmark_is_not_a_machine_surface` RED, verified.

2. LANE 1 — the classifier comment claimed `--gpu-layers all|n` was among the
   flags handled here. It is not: `Commands::Run` carries `gpu` and `no_gpu` and
   nothing else; `--gpu-layers` belongs to `apr serve`. So
   `layers_want_accelerator: false` is CORRECT and the comment was the defect —
   a comment asserting coverage the code does not have. Both comments now say
   the input is genuinely absent for this surface rather than stubbed.

3. LANE 3 — `trace-1789944278.json` (232 lines) was committed to the repo root.
   A runtime byproduct of `test_print_chrome_trace_creates_file`, which writes
   `trace-<epoch>.json` into the CWD when given no `--trace-output`. I swept it
   in with `git add -A` immediately after a quorum round — the exact thing my
   own notes say never to do there. Removed, and this commit stages files by
   name. The test writing into the repo root is a separate defect and is not
   fixed here.

Round 2 was 2 FAIL / 1 NO-VERDICT. Lane 2 has now returned NO-VERDICT twice with
`envelope_status: SUCCESS`, `transport_status: SUCCESS` and non-empty
`raw_bytes` — it answers, and the harness cannot extract a verdict from what it
returns, so width 3 has been delivering two votes.

Verification: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored;
cargo fmt --all --check = 0.

Refs #3602
Pmat-Ticket: PMAT-3602

* roadmap: transcribe #3604 verbatim — round 2's pro lane was right

Criterion (5) had grown an implementation-status clause from PR #3634's
report ('the --json field lands with #3606 StageTimings…') that issue #3604
does not state, and the out-of-scope list was a paraphrase. PMAT-3637 (1)
says every criterion present, none added: the notes now carry the issue's
done_when 1-6, admission rule and out-of-scope list as written.

Refs #3604, #3634
ont-delta: none — roadmap entries only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* audit: quorum-PMAT-3637 — 3/3 PASS, round 3 on the ruled trio

pro / 3.7-flash / 3.6-flash, three conversations, no silences. Rounds 0-2
each said something true about this PR: round 0 judged the wrong ticket;
round 1 found a receipt outside scope and an invented criterion; round 2
found an implementation-status clause in a 'none added' transcription.
None is committed — each judged a diff that no longer exists.

Refs #3604, #3634
ont-delta: none — audit artifact only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3346: quorum verdict 3/3 on e060c7e2c (AD-04)

Round 0 was 1 FAIL / 1 no-verdict / 1 PASS and the FAIL was real (the bias
width, fixed in e060c7e2c). Round 1 on the fixed head: 3/3 PASS,
gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(audits): AD-04 quorum receipt for PMAT-3577 — AGREED 3/3

Three independent agy lanes on distinct models, none in the author's family
(author measured as Opus 5 / claude):

  gemini-3.1-pro-high  = PASS
  gemini-3.7-flash-high = PASS
  gemini-3.6-flash-high = PASS

Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG rather than by editing
the shared config, which three sessions were launching against concurrently.

THE MERGE DID NOT CHANGE WHAT WAS JUDGED. The lanes ran against 7f82e2cd3; the
PR head is 4fbd2ec31, a merge of main. The trees differ by four files, but all
four arrived FROM main, and the PR's own contribution relative to main is
byte-identical across the merge:

  files, pre-merge  : 49        files, post-merge : 49
  only in post      : (none)    only in pre       : (none)
  sha256(diff vs merge-base), pre-merge  : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35
  sha256(diff vs merge-base), post-merge : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35

So this is the same head for review purposes and does not need a new round. A
naive `git diff origin/main..HEAD | sha256sum` from my local head disagrees only
because that head carries this receipt commit, which the lanes never saw —
comparing the receipt's parent is what makes the two comparable.

Refs #3600
Pmat-Ticket: PMAT-3577

* fix(tests): the shape-count ratchet was red — #3605 adds a 6th shape (quorum round 1)

`the_tracked_repo_graph_is_fresh` asserts `shapes_n == 5` and names the five.
`refusal-receipt-v1` makes six, so this PR shipped a CI red that a quorum lane
found and I did not.

MEASURED, not guessed: `pv extract contracts --check` on this branch reports
`shapes_n: 6, triples: 15620`.

The hardcoded count stays hardcoded. A shape added without anyone noticing is
exactly what this assertion exists to prevent, so adding one is SUPPOSED to turn
it red and make you name the new shape. Noted in the comment that the counter is
shared across branches — a sibling PR adding a shape (#3600's parity-receipt-v2)
will need it raised again at merge, which is the ratchet working rather than a
conflict to route around.

Refs #3605
Pmat-Ticket: PMAT-3605

* chore(audits): AD-04 quorum receipt for PMAT-3602 — AGREED 3/3

gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS.
Author measured as Opus 5 (claude); no lane in the author's family. Lane models
set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared
config that other sessions were launching against.

This trio was chosen on measured odds after flash-class lanes returned
NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my
earlier runs). All 15 lanes voted in this batch.

Refs #3638
Pmat-Ticket: PMAT-3602

* chore(audits): AD-04 quorum receipt for PMAT-3605 — AGREED 3/3

gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS.
Author measured as Opus 5 (claude); no lane in the author's family. Lane models
set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared
config that other sessions were launching against.

This trio was chosen on measured odds after flash-class lanes returned
NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my
earlier runs). All 15 lanes voted in this batch.

Refs #3613
Pmat-Ticket: PMAT-3605

* PMAT-3351 (adoption): re-id the fragment — PMAT-3347 is #3348's row on main

Both #3348 (merged; bound the qwen35-e2e equations, closed #3347) and this PR
(the L2 column reads a link, not an index) address issue #3347, and both
claimed roadmap id PMAT-3347. Two PRs cannot share a row: main's PMAT-3347
stays as #3348's, this PR's fragment becomes PMAT-3351 (github_issue 3347,
same title, same notes), and the aggregate is rebuilt from main's roadmap.yaml
plus this branch's fragments — 935 + 1 = 936, nothing re-serialised.

Refs #3347, #3348

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3351: notes transcribe issue #3347's Ask verbatim (bullets 2 and 3; bullet 1 was #3348)

The quorum lanes read pmat work status, i.e. this fragment's notes, not the
GitHub issue. Empty notes would have them judge against a title. The issue
has no done_when section, so its Ask and its two defect statements are
transcribed as written, with the out-of-scope bullet named and no count
offered as a criterion.

Refs #3347

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3338: register the second ticket this PR closes, so the quorum judges the truncate fix as in scope

Quorum round 0 on a4b76397f was 2 FAIL / 1 PASS, both FAILs on one point: the
branch's first commit fixes #3338 (--table panicked on a byte-index cut inside
a multi-byte char) and PMAT-3351 does not ask for it. Both lanes called the
in-scope work correct. The fix is a prerequisite — the L2 column is exercised
through --table on the real corpus, which panicked before it — and #3338 is
an OPEN issue this PR genuinely closes. So it gets its row, transcribed from
the issue, and round 1 runs with --ticket PMAT-3351,PMAT-3338.

Refs #3338, #3347

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3604: the receipt's temp file is private to its writer — a shared name made the atomicity claim false

AD-04 quorum round 0 on #3634, lane 1 (gemini-3.1-pro-high), cited
f2_receipt.rs:180: write_receipt used one shared <sha>.json.tmp, so two apr
runs validating the same model at once could have writer B truncate the file
writer A was about to rename, and A rename B's partial into place. The reader
would classify it Unreadable and validate — the safe direction, done_when 4 —
but the docstring said 'never a truncated one', and that was false.

The temp name now carries the pid and a per-process counter; a failed rename
removes its own temp. A racing test spawns two writers on one path forty
times and parses the survivor each time: it is always one writer's WHOLE
receipt. 14 tests.

Lane 1's other two findings, for the record: --revalidate is done_when 2
verbatim (the fragment carries it), not scope creep; and F2Outcome's fields
are the values the stderr line already prints and done_when 5's data path
for #3606's JSON, not dead code. Two lanes passed the same diff.

Refs #3604

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3604: quorum verdict 3/3 on dd546592b (AD-04)

Round 0 was 2 PASS / 1 FAIL; the FAIL was the shared temp-file race, fixed in
dd546592b. Round 1 on the fixed head: 3/3 PASS, gemini-3.1-pro-high /
pro-low / 3.6-flash-high, each measured, no dissent.

The PMAT-3604 roadmap row was on this checkout UNCOMMITTED for the resolver
(pmat work status reads the checkout); it lands via #3637. The lanes judged
the committed diff origin/main...HEAD.

Refs #3604, #3637

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3351 + PMAT-3338: quorum verdict 3/3 on d915b8f16 (AD-04)

Round 0 (--ticket PMAT-3351 alone) was 2 FAIL / 1 PASS, both FAILs on the
#3338 truncate fix being out of scope; both called the in-scope work correct.
Round 1 with both tickets registered and named: 3/3 PASS, gemini-3.1-pro-high
/ pro-low / 3.6-flash-high, each measured, no dissent.

Refs #3347, #3338

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(contracts): the v1 contract had no equations, and its own denominator guard was never wired

Two reds on this PR, both mine, and the second is the better finding.

1. `contract_data_integrity` is a SHRINK-ONLY ratchet at 444 and this branch made
   it 445. The 445th is `parity-receipt-v1: no equations`. Measured which one
   rather than guessed — `parity-receipt-v2` is not in the list.

   `kind: pattern` correctly declares no SHAPE (a shape over a class nothing
   instantiates passes vacuously, which is #3610's whole lesson), but equations
   are not shapes, and this contract does assert things: PRC1-INV-001 and
   PRC1-INV-002 already said them. They are now written as the two equations the
   integrity check reads — `no_instances` and `refused_by_name` — rather than as
   new claims invented to satisfy a counter.

2. `check_guards_are_wired.sh`: `NEW: parity_receipt_denominator.sh`. I ADDED A
   GUARD IN THIS PR AND WIRED IT INTO NOTHING — the second, independent reader of
   the receipt corpus, shipped where no workflow names it.

   Minutes before finding this I wrote into that same contract's equations, as a
   precondition: "both readers are wired into CI — a refusal nothing runs is not
   a refusal." My own precondition, violated by my own PR, in the same file.
   `check_guards_are_wired.sh` caught what I had just finished writing down.

   Now named in `ci.yml` beside the other cargo-free, model-free case tables. Its
   4-case self-test passes: a legacy record refused by name, a receipt added
   without bumping the denominator disagreeing, bumping it making them agree.

NOTE — THIS TOUCHES `.github/workflows/ci.yml`, which CLAUDE.md lists as a
check-in item. Additive only: one step in an existing guard block, no matrix,
trigger or gate logic changed. Wiring an unwired guard is the minimum the failing
check asks for, and leaving it unwired to avoid the file would be keeping a
refusal that never runs.

VERIFICATION
  cargo test -p aprender-contracts --test validate_contracts contract_data_integrity   1 passed
  bash scripts/parity_receipt_denominator.sh --self-test                               rc 0, 4/4
  bash scripts/check_guards_are_wired.sh                                               PASS (4 -> 3)
  pv validate contracts/parity-receipt-v1.yaml                                         0 errors, valid

DIFF MOVEMENT, for the receipt: the judged diff DID change — an equations block
and one CI step. No behaviour the lanes reviewed was altered; the extractor, the
shapes and the denominator script are byte-identical. AD-04 re-review is the
inspector's call.

Refs #3577
Pmat-Ticket: PMAT-3577

* PMAT-3401 (adoption): the three derivatives 24 new contracts oblige — census, graph, README — regenerated together

This PR adds 24 contracts and regenerated none of the tracked artifacts
derived from the corpus. On #3581 that omission surfaced one per CI round,
each masked by the one before it. All three here, at once, with a pv built
from this tree under a pinned target dir:

  contracts/census.json   1800 -> 1824  (+24, the contracts added)
  contracts/contracts.nt  15,600 -> 15,696 triples  (+96 = 24 x 4; GREW —
                          a drop is the tell for a malformed input or a stale
                          binary; binding.yaml still parses, 156 entries)
  README CONTRACT_COUNT   2 blocks -> 1824 via make readme-sync

The merge took main's generated README blocks over the branch's hand-typed
1866, then readme-sync wrote the measured 1824; the branch's number described
a tree that never existed on main.

Verified: all 24 contracts pv-validate under the pinned binary (control:
main's crux-A-01 valid under the same one); lint_passes_on_real_contracts
green, so the sigma prose ratchet holds; 1666 engine tests; ont4b shapes gate
11/11; test-binding ratchet; readme-sync-check; FALSIFY-README-002; roadmap
additive added=26 deleted=0, aggregate idempotent, 26 fragments present.

Refs #3401

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(run): extract reconcile_and_emit — run() hit cognitive 27 against a 25 ceiling

`check_complexity_ratchet.sh` went RED with:

  RED  NEW  crates/apr-cli/src/commands/run_entry.rs::run  cyclomatic 21 cognitive 27
            (over a threshold, absent from the comparand)

The reconcile-then-emit block I added in this PR is what took `run` over. Moved
into `reconcile_and_emit`, which also gives the ordering decision — machine
surfaces emit on a refusal, the human surface does not — a place to be documented
that is not the middle of a 30-argument entry point.

Now: PASS (D2): e6f77c98c vs 19750b384 measured by pmat 3.41.1 — none new, none grown.

NOTE FOR THE NEXT PERSON: I first re-ran the ratchet against my WORKING TREE and
saw the identical cognitive 27, and briefly concluded the extraction had not
helped. It had. The script measures two REVISIONS — it prints
`merge HEAD <sha>` — so an uncommitted fix is invisible to it. Commit, then
measure.

Tests: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored;
fmt --check = 0; clippy -D warnings clean.

Refs #3602
Pmat-Ticket: PMAT-3602

* PMAT-3496: a registration row for the registration PR — the author's earlier receipt named this ticket but no row existed on the branch

* PMAT-3496: clause (1) transcribes issue #3495's title verbatim (132 Kani harnesses; Verus pilot), not a paraphrase

* fix(guard): a step NAME read as a bare guard_tree.sh dispatch hid four dark guards; the #3305 guard runs nowhere and dies on intel's python 3.10 (#3644)

Refs #3644 #3646 #3305 #3626

THE FINDING, proven by mutation before it was argued. check_guards_are_wired.sh
counts a guard wired-by-dispatch when a workflow runs guard_tree.sh in a mode
whose --dry-run RUN set contains it. Its invocation regex accepts `guard_tree.sh`
followed by `'` (for a quoted run: string). ci.yml has two step NAMES,
`"guard_tree.sh's own case table …"` and `"guard_tree.sh's parallel dispatcher
…"`; the apostrophe matched, neither line says --no-cargo, so each was read as a
bare dispatch in mode "all" -- 55 cargo-classified guards counted wired that
`guard-tree` (--no-cargo) never runs and `guard-cargo` never names. Rewriting
those two names (nothing else) turns the meta-guard RED: 3 -> 7 unwired --
check_model_ladder.sh, check_pathonly_devdeps_unused_in_src.sh,
check_pr_review_counts.sh, check_pr_review_receipt.sh. Three of the four are
cargo-classified by a COMMENT that says the build tool's name; none of those
three invokes it.

check_pathonly_devdeps_unused_in_src.sh (#3305: src/ may not use a dev-dep that
publishing deletes; clean-room red 8/8 across two releases) has therefore run
nowhere since it was written on 2026-09-15. And it could not have run on the
intel hosts: `import os, re, sys, tomllib` at module level, python 3.10.12
there with neither tomllib nor tomli (measured on mac-server tonight; the
sovereign-ci container has no python3 at all). Reproduced with python3.10 -S:
the traceback is swallowed by `out=$(scan … || true)` and the case table prints
"FAIL: missed hit_use" ×4 -- the regex reported broken by an interpreter that
never ran it. Same class as #3626's guard on intel-clean-room-6.

WHAT CHANGES

scripts/check_guards_are_wired.sh
  - `_not_a_name_line` drops `name:` lines before the invocation match, in
    dispatcher_wired() and the per-guard scan. A step name is documentation
    whatever it contains.
  - rows 8-10: the exact ci.yml fixture (`- name: "guard_tree.sh's own case
    table …"`, no run: line) must leave the dispatched guard unwired; a QUOTED
    run: line still dispatches (the control the `'` exists for); a step named
    after a guard wires nothing. Mutant (drop removed): rows 8 and 10 RED.
  - the ledger's ratchet is now set-aperture, owned by this file.

scripts/lib_baseline_ratchet.sh, scripts/check_baseline_ratchets.sh
  - set-aperture gains a NAME-entry admission for sets of FILES: an added entry
    with no `:` is admitted iff the comparand carries that file (at the entry's
    path, or beside the owning guard -- the ledger names siblings by basename
    and guard_tree.sh reads it so) AND the owning guard changed in the diff.
    A file this branch created is refused (PERF-028's shape). Six rows: predates
    -> admitted; branch WROTE it -> refused; no guard edit -> refused; escaping
    path -> refused; beside the OWNER -> admitted; beside a different owner ->
    refused. Lib mutants (always admit / branch removed / sibling resolution
    removed) each turn rows RED; a separate absolute-path check survived its
    mutant because `git cat-file -e` already refuses those, so it is not there.
  - classify: unwired_guards_baseline.txt -> set-aperture, owner named.

scripts/unwired_guards_baseline.txt  3 -> 6, as APERTURE REVEALS with a reason
  per line (each may only leave): check_model_ladder.sh (T-2 release gate, the
  multiplatform_dogfood class); check_pr_review_counts.sh (RED on main today,
  "6 row(s) disagree" -- #3646); check_pr_review_receipt.sh (takes a receipt
  path; nothing passes one -- #3646). check_pathonly_devdeps_unused_in_src.sh
  is NOT ledgered: see below.

scripts/check_pathonly_devdeps_unused_in_src.sh
  - readers tomllib -> tomli -> ENV rc=2 naming python version and $RUNNER_NAME.
    No purpose-built manifest reader: a second TOML implementation over 78
    manifests, where a wrong "has source" is a silent PASS on exactly the
    #3305 class, is not a fallback worth shipping. UNMEASURED-with-a-name is
    the honest state on a runner with no TOML library.
  - the scanner ends with `SCAN-DONE manifests=N reader=…`; the shell REQUIRES
    it. Output without it (no reader, /bin/false, no python3) is ENV rc=2 --
    never rc=1, never "no violations". The selftest and the tree scan
    propagate 2 instead of swallowing it.
  - five death rows: both readers absent (sys.modules nulled so the imports
    fail for real); interpreter exits 1 silently; no interpreter (127); the
    working interpreter as control; and THE WHOLE GUARD under the intel shape
    (nested, recursion-guarded) -- the call site is where `|| true` hid it.
    Mutants: trailer not required (3 rows RED); trailer printed on the
    no-reader path (1 RED); `|| true` restored (the whole-guard row RED).
  - three comments and one FAIL string reworded so the file no longer says
    the build tool's name with a space after it: it invokes none, and that
    substring is what CARGO_RE classifies on. guard_tree.sh --dry-run
    --no-cargo now lists it as `run:`; bare run on this tree rc=0 in 3 s.

Passes: check_guards_are_wired --self-test 10/10; check_baseline_ratchets
--self-test 61 rows; pathonly --selftest under python 3.13.1 and 3.10.12
(tom…
guyernest pushed a commit to guyernest/aprender that referenced this pull request Sep 29, 2026
…ate turns down, instead of "parity disproven" (5) (paiml#3689)

* fix(pv): --table panicked on the real corpus — it cut a property by BYTE index (#3338)

`pv proof-status contracts/ --table` exits 101 on `contracts/`:

    byte index 40 is not a char boundary; it is inside '∈' (bytes 39..42)
    of `for Q4_K_M Qwen2.5-Coder, quantization ∈ {Q4_K, Q6_K}`

The column width is a byte count (`property.len()`, capped at 40) and
`truncate` sliced `&s[..max]`, so any property whose byte 40 lands inside a
multi-byte char panics. Eight contracts in `contracts/` do; the first one
walked is `apr-inspect-quantization-v1.yaml`.

The budget stays a byte budget — the table is laid out in bytes — and the
cut now walks back to the nearest char boundary.

Two tests, both RED before this commit (each panicked at
obligation_matrix.rs:169): the helper row, and one through
`format_obligation_table` because that is the path the operator hit. The
helper fixture asserts `!s.is_char_boundary(40)` first, so it cannot
silently stop proving anything.

Pmat-Ticket: PMAT-3347
Refs #3347, #3338

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(pv): the L2 column read an INDEX, not a link — so 3,573 of 3,753 obligations ticked (#3347)

`obligation_matrix` computed `l2_tested` as `idx < falsification_tests.len()`.
Obligation 3 was "tested" because the contract happened to hold 4 tests,
whoever those tests were about. All 7 obligations of
`qwen35-e2e-verification-v1` showed ✓ before a single test existed. The
fallback was a substring match between the obligation's `property` prose and a
test's `rule` prose, which is an inference, not a claim.

WHICH LINK, decided by counting the corpus (1,842 files, 3,792 obligations,
4,691 falsification tests), not by preference:

  proof_obligations[].discharged_by   89  (62 `falsification_tests[N]`, 13 a
                                           test id, 1 a YAML sequence, rest
                                           comma lists / prose / a kani id)
  falsification_tests[].binds_to      38
  falsification_tests[].obligation    26  (12 name an obligation id, 6 the
                                           exact property text, 8 dangle)
  proof_obligations[].id             180 of 3,792
  kani_harnesses[].obligation      1,918 — that is the L3 column, not this one

All three L2 spellings are DECLARATIONS by the contract author, so all three
are read. `binds_to` is a serde alias of `obligation`, which is safe only
because no entry carries both keys — checked, because serde turns that into a
`duplicate field` parse error rather than a silent pick.

Not read: `applies_to`. 12 `binds_to` values match one, but `AppliesTo` is an
enum with `#[serde(other)] Other`, so the string is discarded at parse and
there is nothing left to compare. Those 12 report `?`, not a false ✗.

WHAT THE COLUMN NOW SAYS. Three values, because "no test covers this" and
"nothing here says which test covers what" are different facts:

  ✓ Tested   a test in this contract cites this obligation
  ✗ Untested the contract's links resolve and none names it — or it ships
             no falsification test at all, which is a reading, not a gap
  ? Unknown  no readable link; not measured. An unread window is Unknown,
             never a tick

MEASURED over `contracts/`, and the drop IS the point — it is what the old
column was hiding:

              before   after
  L2 ✓         3,573      86
  L2 ✗           180      65
  L2 ?             0   3,602
                       (3,753 obligation rows, 873 contracts)

Nothing was adjusted to keep the number up, and no threshold was added.

The two link fields did not exist on the structs, so both keys were written to
disk and silently dropped on parse — the shape of #3314 (`id`) and #2465
(`test_harness`). `discharged_by` is typed `Citation` (scalar | comma list |
YAML sequence) because `Option<String>` failed the WHOLE corpus on
`publish-manifest-v1`: `invalid type: sequence, expected a string`.

RED first: `l2_does_not_tick_for_an_obligation_no_test_cites` — two
obligations, two tests, both citing OB-A. It asserted through the rendered
table (the surface that was lying, and API-stable), so it is the SAME test
before and after: on the old code `Beta holds | ✓`. Its OB-A arm keeps a fix
that merely stopped ticking everything from passing.

Out of scope_paths, and forced: `lint/strict_test_binding.rs` holds the only
EXHAUSTIVE `FalsificationTest` literal in the tree (no `..Default::default()`),
so no schema field can compile without that one line.

Pmat-Ticket: PMAT-3347
Refs #3347, #3091, #3114

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(pv): single-file --strict-test-binding reported every ref missing — it now refuses (#3347)

The gate resolves cited test names against a source index rooted at the
contract path's PARENT. The directory form gets the repo root and finds
`crates/`; the single-file form gets `contracts/`, which holds no source, so
every cited ref resolves to nothing and all of them are reported missing.

Measured on `contracts/pv-artifact-kinds-v1.yaml`, a control whose eight refs
all resolve:

    pv lint contracts/pv-artifact-kinds-v1.yaml --strict-test-binding
        total_refs 8, existing 0, missing 8
    pv lint contracts/ --strict-test-binding
        total_refs 548, existing 521, missing 27  <- none of the 27 is this one

A control contract failing identically to a broken one is a gate that cannot
discriminate, and it fails SILENTLY: `passed` is true in non-strict mode, so
the run still says `Result: PASS` while printing eight false findings.

REFUSED rather than repaired. The scan root is computed in
`provable_contracts::lint::run_lint`, outside this ticket's scope; a refusal
lives in the caller, is honest, and cannot be mistaken for a clean bill. If
the root is later made explicit (a `LintConfig` field the CLI can feed, which
is what `--crate-dir` does NOT do today), this refusal is what should be
deleted.

Exit 1, not the exit-2 `decline:` class: exit 2 belongs to `ZeroContracts`,
whose message ("0 contracts under ...") would be false here — there IS a
contract; it is the gate that cannot run over it.

Two tests, in the CI-wired `cli_integration` target: the refusal names the
flag and prints no PASS and no findings, and a directory-form control proves
the refusal is specific to the single-FILE form rather than to the flag.

Pmat-Ticket: PMAT-3347
Refs #3347

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(crux): AutoGluon becomes a CRUX competitor — category O, 24 contracts, 25 tickets, registry edit + mutation proof

Research of ../autogluon (1.6.3 @ 77946149) as a competitive-research
source for aprender, landed the way category N landed linfa and burn
(#3169): the competitor is admitted to the CLOSED registry, every story
is a real contract with falsification gates, every story is a registry
row, and every missing/partial story has a GitHub issue and a roadmap
fragment.

What AutoGluon is, measured from the tree rather than recalled: three
predictors. TabularPredictor (65 public methods, 11 presets, 24 model
families of which 8 are tabular foundation models added in 1.4-1.6),
TimeSeriesPredictor (Chronos-2/Toto-2 pretrained, 30+ local/deep models,
16 metrics incl. WQL/MASE/RMSSE, auto backtesting since 1.5) and
MultiModalPredictor. Evidence under evidence/crux/autogluon/.

What aprender has, measured at eb262f8eb: automl/ is a single-estimator
hyperparameter tuner (TPE, grid, random, DE, TimeBudget, EarlyStopping);
time_series/ is one univariate ARIMA; encoders, calibration, SHAP/LIME/
permutation importance and KFold/cross_validate exist as building
blocks. No predictor-level fit(label), no leaderboard, no bagging,
stacking or greedy weighted-ensemble selection, no panel forecasting,
no quantile forecast metrics. The gap is the AutoML UX, not the
algorithms.

Category O — AutoML Parity — 24 stories: 9 P0 (the README hello-world:
fit(label), problem-type inference, presets, leaderboard, feature
pipeline, weighted ensemble, time budget, panel forecaster, quantile
metrics), 9 P1 (bagging, stacking, importance, threshold calibration,
deployment artifact, tabular foundation model, backtesting, local
baselines, pretrained forecaster), 6 P2 (refit_full, distill,
infer_limit, fit diagnostics, memory-aware fit, covariates).
MultiModalPredictor, autogluon.cloud, MLZero and Ray-parallel fits are
CUT on the epic with reasons.

  CRUX_COMPETITORS: [&str; 14] -> [&str; 15]   + autogluon

Not a BEAT pillar: aprender claims no pinned-benchmark win over
AutoGluon. Both registry tests that keep BEAT_INCUMBENTS and
CRUX_COMPETITORS apart are extended, not worked around.

Tickets: epic #3370, stories #3371-#3394, label pareto-autogluon.
Roadmap: 25 fragments under docs/roadmaps/entries/, roadmap.yaml
regenerated by the aggregator (idempotent check passes).

Spec: docs/specifications/crux-competitive-research-ux-workflows.md
v2.2 -> v2.3 — §3 gains rows for linfa+burn (category N, which #3169
never recorded there) and AutoGluon; §5 gains Category O; §6 notes
that coverage_intake in the YAML is the source of truth.
coverage_intake 267 -> 291 (partial 72 -> 77, missing 152 -> 171).

Verification:
- pv built from THIS tree validates 25/25 (24 new + master). The stale
  ~/.cargo/bin/pv rejects crux-O-01 with CRUX-002 — the behavioural
  delta proves the registry edit engaged.
- Mutation-verified: deleting "autogluon" from CRUX_COMPETITORS turns
  competitor_registry_covers_the_corpus_vocabulary RED with "autogluon
  is used by contracts/ and must stay in CRUX_COMPETITORS" and
  the_real_crux_registry_rows_are_all_in_domain RED. Restored: 20/20.
- cargo test -p aprender-contracts --lib: 1526 passed, 0 failed.
- Every falsification gate is LIVE-PENDING prose (no `::`), so
  strict-test-binding has nothing to refuse; the obligations are
  RECORDED as unfalsifiable-by-absence, not satisfied.
- README CONTRACT_COUNT regenerated 1835 -> 1866 by readme_sync.sh.
- Guards: roadmap fragment/ids/sorted/additive/completion, contract
  test-binding and enforcement, shell-lint ratchet, hardcoded paths,
  readme claims, grep -q ratchet — all rc=0.

Pmat-Ticket: PMAT-3370

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(crux): bind the admission PR to its own ticket PMAT-3401 — the epic's acceptance criteria describe the programme, not this diff

Round 2 of the quorum read the epic (PMAT-3370) as the ticket and refused
the admission for not implementing the 24 stories. The admission is its
own unit of work with its own done-when; this fragment says so.

Closes #3401
Pmat-Ticket: PMAT-3401

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* chore(crux): PMAT-3401 title inventories the diff exactly — 26 fragments (its own included) and the scaffold CATEGORY_NAMES edit

Quorum round 3 lane 1 refused on two literal mismatches between the
ticket and the diff: '25 roadmap fragments' (there are 26 once this
ticket's own fragment lands) and an unlisted edit to
scripts/crux_scaffold_contracts.py. The title now lists every path.

Pmat-Ticket: PMAT-3401

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* roadmap(PMAT-3495): VERIFY-001 on 0.71.0 — run the 132 Kani harnesses in CI (proof credit from runs, not declarations), then a Verus pilot on one dequant/parser function (#3495)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* quorum(PMAT-3496 brief, PR #3496): 3/3 PASS on 57e0e4dfb — gemini lanes, measured; author claude-fable-5-1

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* PMAT-3577: the logit-parity receipts under contract — parity-receipt-v2, extract:parity-receipt, and the 7 back-filled

Until this row the logit-parity records under evidence/parity/** had NO validator
of any kind. Not a weak one — none. That is why seven of them sat in the tree
carrying no comparator for months: there was nothing that could have noticed.

WHAT LANDS
  · contracts/parity-receipt-v2.yaml — three shapes: parity-receipt-complete
    (closed, ignoredProperties empty), parity-comparator-self, -oracle. The
    subset has no sh:or, so the comparator split is two shapes over two
    subclasses the extractor assigns by kind.
  · contracts/parity-receipt-v1.yaml — the retired layout, recorded with NO
    shape: all instances were migrated, and a shape whose target class nothing
    instantiates passes vacuously. The enforcement that replaces it fires — the
    extractor refuses an unmigrated record BY NAME, and so does the predicate.
  · ontology/extract/parity_receipt.rs — record / unmigrated / other, and
    skipping is never silent.
  · scripts/parity_receipt_denominator.sh + evidence/parity/EXPECTED_RECEIPTS —
    the count is pinned by an INDEPENDENT predicate. An extractor checked
    against a number the extractor produced proves nothing.
  · Unknown{ExtractorMiss}, exit 2 — a new element of the verdict lattice. An
    extractor that silently saw the wrong corpus reports the same "no
    violations" as one that saw all of it.
  · the 7 records migrated to v2 and back-filled with comparator {kind: self,
    reason}, in this commit, as the row requires.

THE DENOMINATOR IS 7, NOT 8. The #3574 receipt is in PR #3575, still open; it is
not on main. Measured: 113 files under evidence/parity, 7 parity records, 0 with
a comparator. #3575 bumps it to 8 when it lands — this row's own falsifier on its
first real use.

check_parity_receipt.sh IS NOT TOUCHED, and that is the finding, not an omission.
It validates the THROUGHPUT family (instrument, protocol_ref, lanes[],
decode_tok_per_sec, the #2696 cross-class defect); a logit record has never
carried one of those keys. Folding it in — as item 6 asked — would have deleted
the #2696 validator from a family nobody was watching. One validator per artifact
family, and the discriminator is the artifact's required keys, never its filename.

THE BACK-FILL IS A RELABEL. `raw` is the original apr parity --json document key
for key; every envelope field is quoted from committed evidence named in each
record's provenance.record. model_sha256 is carried only by the one record that
measured it: hashing the files today and attaching that to a receipt about
2026-09-06 would be a claim about a different world wearing a witness's clothes.

One derived field was wrong first time and is worth recording: result.verdict
copied raw.parity — apr's own per-position flag — which says PASS for the two
1.5B cells their own RECORD.md calls RED. It now resolves the threshold from
thresholds.yaml and reproduces all seven readings the RECORD.md files state,
both REDs included. No threshold is ever typed into a shape.

NOTHING IS ARMED. armed_shapes lives in lint-baseline.json, a shared file this
row may not touch (decision 7); the shapes are computed and reported, as
ladder-green was at ONT-4c1. Arming is a follow-up with the label.

CONTROLS, both directions: 7 violations with the comparator stripped from all
seven → 0 as committed; exactly 1 for a single plant, naming focus node and
property; widening sh:in to accept `oracle` turns ont4c3_parity_receipts RED, and
mutating only the real contract turns the fixture-drift test RED; an unmigrated
record declines at exit 2 naming the file; 2 receipts against a denominator of 1
declines naming both numbers; pass / fail / decline are 0 / 1 / 2.

  cargo test -p aprender-contracts --lib ontology::extract::parity_receipt  10 ok
  cargo test -p aprender-contracts-cli --test ont4c3_parity_receipts         8 ok
  bash scripts/parity_receipt_denominator.sh --self-test                     4 ok
  pv lint contracts --gate {sigma,relations,shapes}          Pass, 0 violations
  pv extract contracts --check                                            rc 0

ONT-4c3 is NOT bound in the ONT-001 ledger: that ledger is paiml/infra's
paiml-ontology.md, where v4.8 still defines ONT-4c3 as KERNEL receipts. The
re-scope is an infra PR, not this one — raised with the cop rather than left as a
checked box. Receipt: docs/audits/impl-PMAT-3577-receipt.md.

Refs #3577, #3576, #3269, #3575, #3567

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: name it WrongCorpus, not ExtractorMiss — a prefix of its opposite is a defect

`ExtractorMissing` already meant the opposite thing: an extractor that does not
exist. `ExtractorMiss` would have sat beside it in the same 18-element lattice,
one letter apart, with the shorter a PREFIX of the longer — `grep ExtractorMiss`
matches both, and any substring test over the reasons merges them silently. That
is the defect class this repo keeps paying for (#3573 today), so the name now says
what happened: the extractor RAN and read a corpus the tree does not declare.

Renamed before it reached main, when a rename is a sed and not a migration.
Caught in review by the cop.

Refs #3577, #3573

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: the ONT-4c3 divergence is filed as paiml/infra#814, not left in a thread

Two repos holding different definitions of one identifier is a row, not a flag in
a message: nothing collides where a tool would see it, so it collides in a
person's head months later when they implement the ledger's meaning and find
their correct work unusable.

Refs #3577, paiml/infra#814

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* #3605: removed_by gets a shape — refusal-receipt-v1, with a closed sentinel set so the required field manufactures nothing

`removed_by` had 0 occurrences tree-wide, no schema and no validator. Every
refusal #3597 writes would have minted an unenforced convention, and a tree full
of consistent-looking `removed_by:` lines reads as validated when it is
decoration.

PHASE 0 — ITS OWN CONTRACT, AND NOT BECAUSE IT IS EASIER TO WRITE

Option B is refused on a fact: `parity-receipt-v1` DOES NOT EXIST AT HEAD. It is
in unmerged #3600, and there it is deliberately the RETIRED layout with no shape
and no instances. A live validated field on a superseded contract with zero focus
nodes is a field nothing can carry. Independently, by the discriminator this tree
has now used twice — an artifact family is named by its REQUIRED KEYS, never by
its filename — a parity receipt requires host, backend, comparator,
threshold_source and per-position metrics; a refusal requires a verb, a reason,
an exit code and removed_by. They share no required key.

A THIRD option was considered and refused, and it is the more attractive one:
put removed_by on apr-cli-commands-v1.yaml, where the verb already IS the focus
node and the universe is already the trustworthy 111. Refused because the
registry is the UNIVERSE and #3597's method is to DIFF the registry against the
buckets. If the buckets live in the registry, the denominator and the numerator
are the same artifact and the diff is vacuous by construction. The registry says
what EXISTS; a refusal says what was TRIED. Keeping them apart is what lets that
diff be a real diff.

THE FORCED-BINDING TRAP, AND WHY minCount 1 IS SAFE HERE

A required field with no escape manufactures false data: `pv validate` requires
kani_harnesses so authors fabricate one, and apex#57 had lean_theorem copied
verbatim into twelve contracts resolving to nothing, gate green throughout. The
escape is a closed sentinel set, so "there is legitimately nothing here" is
SAYABLE and still CHECKABLE:

  v<major>.<minor>   the release that removes it
  never              refused permanently by design
  unscheduled        a defect with no release chosen

tbd, soon, n/a, pending, a bare `0.70`, a patch-level `v0.70.1` and a git sha are
all RED. A version rather than a sha because a refusal answers "which release do
I need?" and a sha is precise about the tree and silent about the boundary; the
sha form is made INVALID rather than discouraged, because a shape that permits
two spellings gets both.

BOTH ARMS, and the second is the point: 9 cases over 8 fixtures. refusal-ok
carries all three declared forms and passes with 3 focus nodes;
refusal-undeclared-sentinel plants the PLAUSIBLE `tbd` and must go red. Mutation
proved red-capable: widening the pattern to accept tbd fails exactly
a_plausible_but_undeclared_sentinel_is_refused and
accept_and_refuse_are_distinct_answers, and restoring makes 9/9 green. Fixtures
carry the real contract byte for byte, so widening it without them is caught too.

No new extractor: entity {type: json, ref} + vocabulary is the existing machinery
for "validate this document", and a second reader for one more family is the
thing this tree keeps filing against.

done_when 6 — the interim re-check list is EMPTY: `removed_by` still has 0
occurrences at HEAD, so no refusal written under the interim needs revisiting.

The ledger is seeded with the one refusal already measured (`apr bench` refuses
qwen35) so the shape has a real focus node instead of passing vacuously over an
empty list. WHICH verbs land there is #3597's bucket, not this row's.

Refs #3605, #3597, #3600, #3080

Pmat-Ticket: PMAT-3605

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: count parity-receipt in by_entity_type — ONT-001 v4.10's probe reads ABSENT otherwise

The spec's probe asks `by_entity_type["parity-receipt"] == 7`. Measured on this
branch before the change: the map carried pv-contract, gguf, apr-model, code and
lean, and NO parity-receipt key. The seven records were there — by_shape showed
parity-receipt-complete=7 — but the entity-type map did not carry them, so the
probe would have read ABSENT.

AN ABSENT KEY IS NOT ZERO. A consumer treating it as one measures nothing and
calls it a pass — the same shape as #3610, one map over.

Registering the entity type in Sigma was not enough; it has to be counted where
the probe looks. A test now asserts the key exists AND that it equals the shape's
own focus-node count, so the two numbers cannot drift apart.

Refs #3577, #3610, paiml/infra#831

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3577: regenerate contracts.nt with a pv built from THIS worktree — the shared target dir served another branch's binary

The merge commit's contracts.nt was 107 triples short: every parity-receipt and
parity-comparator node was missing, because `cargo metadata`'s target_directory
is shared across worktrees and the pv on PATH had been built from a different
branch minutes earlier.

AND `pv extract contracts --check` PASSED ON IT, because the check re-derives the
graph with the same binary. A stale tool comparing an artifact against its own
re-derivation agrees with itself about nothing being there — the derived file and
the checker were wrong in the same direction, which is the only way that gate can
fail to fire.

Rebuilt with CARGO_TARGET_DIR pinned to this worktree; the 107 triples return and
by_entity_type[parity-receipt] reads 7 rather than absent.

Refs #3577

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* register the new CLI test in scripts/tree_reader_tests.txt

The registry is derived from the tree and drifts the moment a test that reads the
tree is added without listing it. Each of the three new tests failed this guard on
its OWN branch — not a shared commit, as first read.

Pmat-Ticket: PMAT-3598

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* register the new CLI test in scripts/tree_reader_tests.txt

The registry is derived from the tree and drifts the moment a test that reads the
tree is added without listing it. Each of the three new tests failed this guard on
its OWN branch — not a shared commit, as first read.

Pmat-Ticket: PMAT-3598

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* #3600: the migration broke TWO readers of the records, including the release judge

I asked "who else reads this?" for the <unk> assumption and did not ask it for my
own migration. Two consumers read the legacy top-level metrics:

  crates/apr-cli/src/commands/parity_admission.rs  -> mac-check RED
  scripts/check_model_parity.sh                    -> C14, the RELEASE judge

The second is the one that matters: C14 is the pre-publish dogfood's parity gate,
and on the migrated records it reported "no per-position metrics in the output"
for a record that is fine. The migration would have taken the release gate down.

THE RULE, STATED ONCE IN BOTH READERS: a v2 receipt EMBEDS the raw
`apr parity --json` document under `raw`, so look inside the envelope when there
is one. A fresh `apr parity` run is the raw document itself and carries the
readings at the top level. Those are two different INPUTS — a tool's output and an
archived receipt quoting it — not two spellings of one, which is the distinction
that makes this a rule rather than the permissiveness #3613 refuses.

The self-test's fixture builder read the same way, so the must-RED twin was being
built from a KeyError and that control could not have fired.

Verified: 7B sentinel PASS, 1.5B sentinel RED (unchanged from pre-migration),
check_model_parity.sh --self-test 26/26, parity_admission 19 tests.

Refs #3600, #3577

Pmat-Ticket: PMAT-3577

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3604: the F2 hybrid guard runs once per (model sha256, apr version, device) and leaves a receipt

f2_validate_qwen35 proves the CUDA hybrid forward against a 64-position CPU
reference before it will serve a token, on EVERY apr run. Measured (#3598
row 1): 67 % of a 14 s time-to-first-token, 90-93 % of it the CPU forward.
The guard is right to exist and wrong to run per call.

Now: the first run of a (model sha256, apr version, device) triple validates
and writes a receipt; a later run whose triple matches reads it and skips the
forward; `apr run --revalidate` forces a fresh run and rewrites it.

MEASURED HERE, RTX 4090, Qwen3.5-4B-Q4_K_M, the 144-word row, --max-tokens 1,
GPU occupancy recorded before every run, binary built from this tree under a
PINNED target dir (the shared one handed me another worktree's binary first):

  --revalidate, warm cache   17.18 s wall   guard 9,747 ms on 65 positions
  receipt hit, warm cache     7.39 s wall   guard 0 ms      sha256 1,173 ms

  -9.8 s wall; guard 9,747 -> 0; the key costs 1.17 s/run (2.74 GB at
  2.3 GB/s) and is printed separately so it cannot hide in either number.

THE RECEIPT IS THE VALIDATION, WHICH IS WHY IT IS STRICT. Every path that is
not "three keys match" validates, and the three planted-receipt falsifiers
were run END TO END in the real binary, not only as unit tests:

  wrong model sha256   -> re-validated: "receipt is for model bbbb…, this file is 00fe…"
  wrong apr version    -> re-validated: "receipt written by apr 0.61.0, this is 0.68.2"
  wrong device         -> re-validated: "receipt written for NVIDIA GB10, this device is …4090"
  corrupt file         -> re-validated: "receipt unreadable (…: not a receipt)"
  missing file         -> re-validated: "no receipt for this model"

Absence is never consent, and absence and unreadability are told apart.

A RECEIPT IS WRITTEN ONLY AFTER A VALIDATION THAT JUDGED SOMETHING. The guard
has three early exits that let the GPU serve without comparing a position —
SKIP_PARITY_GATE=1, a probe under two tokens, a CPU reference that would not
run. f2_validate_qwen35 now returns F2Verdict {Accepted{positions_judged},
Rejected, NotJudged} instead of bool, and only Accepted writes; otherwise a
one-token prompt would "validate" the triple for every prompt after it.

The decision table (f2_receipt.rs) is pure and CUDA-free, so its 13 tests run
on every build. --revalidate reaches the guard through the same env seam the
guard already reads SKIP_PARITY_GATE from, rather than a 38th positional
parameter on run_entry::run and six forward signatures #3606 is changing.

done_when 5 is partial and says so: [source=receipt|fresh] is on the guard's
stderr line; the `apr run --json` field lands with #3606's StageTimings, and
F2Outcome{source, validate_ms, sha256_ms, receipt_path} is returned to the call
site for exactly that. #3606 and this PR both edit f2_validate_qwen35's return
path; whichever lands second reconciles ~10 lines, and #3606's lane is told.

Evidence: evidence/perf/3604/MEASUREMENT.md + the stderr of all nine runs.

Refs #3604, #3596, #3598, #3606

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* roadmap: mint PMAT-3604 so the AD-04 quorum for #3634 can run

pmat work status PMAT-3604 refused ('Item not found'): #3604 was minted as a
GitHub issue with a done_when but never as a roadmap entry, and the quorum
script hard-requires the work item. Fragment + aggregate, nothing else.

Refs #3604, #3634
ont-delta: none — a roadmap entry; no ontology entity, shape, reason or resolution

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(run): --gpu that fell back to CPU reported success — the refusal had no caller (#3602)

`apr run --gpu` on a model whose GPU attempt is rejected at runtime printed a
result and exited 0. Measured on lambda (RTX 4090 sm_89, apr 0.68.2 e6f77c98c,
qwen2.5-coder-0.5b-instruct-q4_k_m):

  exit=0
  stderr: warning: GPU output diverges from CPU at position 1 (cosine 0.4153)
  stdout: { "used_gpu": false, "inference_time_ms": 33646.13 }

33.6 seconds on the CPU, reported as a successful --gpu run, with `used_gpu:
false` as the only signal — the same value a deliberate CPU run reports.

The decision was already made, recorded and unit-tested. `registry::after_generation`
implements it (R-0b, #3002/#3042): a FORCED accelerator that fell to CPU is a
refusal, a DEFAULT selection that fell to CPU gets a corrective line. A git grep
found its only callers were its own tests. `registry::announce` and
`registry::parity_line` are in the same state, which is why a real run prints
zero `selected:` and zero `parity:` lines.

So this is not a new policy. It is the recorded one, reached from `apr run`:

- `dispatch.rs` classifies the request ONCE via `registry::Request::wanted()`
  rather than re-deriving "forced" — two spellings of one rule drift apart.
- `run_entry::run` calls `reconcile_accelerator` before any success output.
- `--json` gains a `backend` object: `requested` / `ran` / `fell_back`, because
  `used_gpu: false` alone collapses "ran on CPU deliberately" with "asked for
  the GPU and was refused it".

`accel::tests::every_accelerator_surface_calls_the_refusal` passed throughout:
it guards `ensure_available`, the BUILD-time refusal, not the RUNTIME one. A
guard over one of two refusals reads as coverage of both.

Falsifier, both directions (`run_tests_accel_reconcile.rs`, 6 cases). Planting
the pre-fix behaviour (`reconcile_accelerator` → `Ok(None)`) turns
`a_forced_accelerator_that_ran_on_cpu_is_refused` and
`the_json_distinguishes_a_deliberate_cpu_run_from_a_rejected_gpu_run` RED while
`a_forced_accelerator_that_actually_ran_on_gpu_says_nothing` stays green — so
the tests discriminate rather than just failing. `used_gpu: None` is Unknown,
not a fallback: refusing on it would make every non-reporting path a hard error.

NOT included, deliberately: the rejection's REASON (cosine, position) is on
stderr but not in the JSON. It is produced inside realizar's F2 gate and no
channel carries it to the CLI; adding one is a #3606-shaped follow-up, not
something to approximate with a guess here.

Refs #3602, #3483
Pmat-Ticket: PMAT-3602

* test(PMAT-3346): the measured Qwen3.5-0.8B inventory the dense path cannot reproduce

QE2E-INV-001 could not be judged because nothing in the tree held a MEASURED
Qwen3.5 tensor inventory to judge against. This adds one: the 320 tensors of
~/models/Qwen3.5-0.8B-Q4_K_M.gguf (sha256 bd258782...dc517), read straight from
the GGUF header rather than from a model card.

It already falsifies the current arithmetic. Dense/GQA accounting applied to
that file gives 644,400,128 against a measured 752,393,024 — short by
107,992,896, 14.4% of the model, because 18 of the 24 layers are Gated DeltaNet
and no term here counts their conv, gate, state or output projections.

Two shapes in the file are also not what dense accounting predicts, and both are
pinned: attn_q is [1024, 4096] = 2 * num_heads * head_dim (the q projection
emits the attention output gate alongside the query; attn_output [2048, 1024]
confirms num_heads * head_dim = 2048), and attn_q_norm/attn_k_norm are present
at head_dim. The file is also TIED — it has no output.weight.

Refs #3346

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(PMAT-3346): ModelConstraints carries the gated-DeltaNet shape, so a hybrid model can be counted

contracts/model-families/qwen3_5.yaml declares inner_size, state_size,
conv_kernel, group_count and full_attention_interval under constraints:, and
ModelConstraints carried none of them. A Gated DeltaNet layer's parameters live
entirely in those dimensions, so every consumer of the descriptor counted
Qwen3.5 as if three quarters of its layers did not exist.

Carried through as ModelConstraints::deltanet: Option<DeltaNetShape> — the
runtime YAML loader (parsing.rs) and the compiled-in registry (build_parsing.rs
+ build_codegen.rs) both populate it, and FALSIFY-MF-QWEN35-010 pins that the
declared values survive the trip and that no other family acquires a shape it
never declared. qwen3_5.yaml is the only descriptor with these keys, so every
other family keeps byte-identical accounting.

model_arithmetic gains gated_deltanet_layer_params (one term per GGUF tensor:
attn_qkv, attn_gate, ssm_conv1d, ssm_alpha/beta, ssm_a, ssm_dt.bias, ssm_norm,
ssm_out) and hybrid_layers (the interleaved schedule). attention_layer_params
gained two terms the real file has and dense accounting did not model: the
gated q projection (2*n_h*d_k) and the q/k norm vectors.

Falsified against a real model, not against itself: fed the 0.8B configuration,
the equation now reproduces the 320-tensor inventory of
Qwen3.5-0.8B-Q4_K_M.gguf EXACTLY — 752,393,024, both layer kinds matching
tensor for tensor.

QE2E-INV-001 is still NOT asserted, and no range was widened to make it pass.
The 9b descriptor now yields 8,344,907,136, up from 8,208,519,168 but still
0.655B below [9.0B, 9.2B]. The remaining gap looks like descriptor drift rather
than missing arithmetic: 9b keeps inner_size 2048 — the value the 0.8B uses at
hidden_dim 1024 — while quadrupling hidden_dim, and its group_count 8 fails
8 * 128 == 2048, a consistency the measured 0.8B satisfies at 16 * 128. Only a
real Qwen3.5-9B file can settle it; none is on this box.

Refs #3346, #3347

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* contracts(binding): the QE2E-INV-001 note blamed a gap that is now closed

The note said the obligation was undischarged because ModelConstraints does not carry the DeltaNet shape keys. It does now, and the 0.8B count reproduces the real GGUF exactly. What actually blocks the obligation is descriptor drift at the 9b variant.

Pmat-Ticket: PMAT-3346
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3346 (adoption): the QE2E-INV-001 note lost its closing quote, and pv extract shrank the graph by 356 triples instead of refusing

d3cc76f1f rewrote the note on the QE2E-INV-001 binding and dropped the
trailing `"` — contracts/binding.yaml stopped being valid YAML at line 764
(`found unexpected end of stream`). Nothing in the PR noticed because
`pv extract contracts` does not refuse a binding registry that will not
parse: it emitted a graph with 15,244 triples where main has 15,600 —
every bound symbol AFTER the broken entry (prune::run, distill::run,
harness_ir::*, ptx_explain::run, …) silently gone — and `--check`
would have agreed with itself. Found while regenerating the derivative
for this adoption, by the drop, not by any gate.

One character. With it, binding.yaml parses (156 entries, same as main)
and the extraction is byte-identical to main's committed contracts.nt, so
this PR owes no graph change after all.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(roadmap): mint PMAT-3602 as a work item so the AD-04 quorum can resolve it

`quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work`
reads the aggregated roadmap rather than an id counter. The branch carried the
trailer `Pmat-Ticket: PMAT-3602` while no such work item existed — a GitHub
issue number is not a work item.

The id is DERIVED from the issue (`pmat work add --github-issue` calls that
"the path to prefer": GitHub allocates centrally, so two agents cannot be handed
the same id, which `max(id)+1` cannot promise). Fragment, not a direct edit to
roadmap.yaml: additive, one file, and it does not make every stacked branch
dirty on the same lines.

Acceptance is hand-entered from the issue's done_when — `pmat work add` derives
neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414).

Guards: diff_additive, fragment_required, ids_unique, sorted,
completion_is_cited all exit 0; `make roadmap-aggregate-check` reports
`roadmap.yaml == aggregate(48 fragment(s)), idempotent`.

Refs #3602
Pmat-Ticket: PMAT-3602

* roadmap: PMAT-3637 — the registration row this PR is, so its quorum judges fidelity not implementation

Round 0 ran with --ticket PMAT-3604 and all three lanes FAILed for the same
reason: the diff implements none of PMAT-3604's criteria. Correct — it was
never meant to; it registers the ticket so #3634's quorum can start. The
receipt is kept as quorum-PMAT-3637-r0-misframed-as-3604.json (a FAIL is a
record, not a mistake to erase). PMAT-3637 is the row this diff satisfies:
fidelity of the PMAT-3604 entry to issue #3604's done_when, additive
aggregate, nothing outside docs/roadmaps/. Round 1 runs against it.

Refs #3604, #3634
ont-delta: none — roadmap entries only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(roadmap): mint PMAT-3605 as a work item so the AD-04 quorum can resolve it

`quorum-review.sh` refuses without `pmat work status <id>`, and `pmat work`
reads the aggregated roadmap rather than an id counter. This branch carried the
trailer while no such work item existed — a GitHub issue number is not a work
item until something mints it.

The id is DERIVED from the issue: `pmat work add --github-issue` calls that
"the path to prefer", because GitHub allocates centrally and `max(id)+1` cannot
promise two agents different ids. A fragment rather than a direct roadmap.yaml
edit — additive, one file, and it does not make every stacked branch dirty on
the same lines.

Acceptance hand-entered from the issue's done_when; `pmat work add` derives
neither `spec:` nor `acceptance_criteria:` (paiml-mcp-agent-toolkit#1414).

Guards: diff_additive, fragment_required, ids_unique, sorted,
completion_is_cited all exit 0; `make roadmap-aggregate-check` idempotent.

Refs #3605
Pmat-Ticket: PMAT-3605

* PMAT-3346: the roadmap row, so the AD-04 quorum for #3350 has a ticket to judge against

Acceptance transcribed from issue #3346 as this PR answers it (the type that
reads the descriptor was wrong, not the range or the descriptor), with the 9B
range instantiation explicitly out of scope until a real 9B GGUF exists.

Refs #3346

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* roadmap: round 1 said two true things — fix both

Lanes 1+2 (gemini-3.1-pro-high, gemini-3.8-flash-high) FAILed PMAT-3637 on
(a) the round-0 receipt committed under docs/audits/, which criterion (3)
forbade — dropped; the receipt judged the wrong question and its ticket
field would read as a verdict on PMAT-3604; and (b) a criterion in the
PMAT-3604 notes that issue #3604 does not state ('Only an Accepted verdict
writes a receipt') — it is PR #3634's design decision, not a done_when
item; removed from the transcription. Criterion (3) now admits this row's
own receipt at docs/audits/quorum-PMAT-3637*.json, which it must, or the
PASS receipt could never be committed.

Refs #3604, #3634
ont-delta: none — roadmap entries only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(run): quorum round 1 found the --json surface skipped and a branch that could never run

Two findings from lane 1 (gemini-3.1-pro-high) of the AD-04 quorum, both
correct, both fixed. Lane 3 passed the PR without seeing either.

1. `reconcile_accelerator(...)?` ran BEFORE `print_run_output`, so a rejected
   `--gpu` run early-returned and `--json` emitted nothing at all — the exact
   surface #3602 item 1 names. Machine surfaces (`--json`, `--stream`) now emit
   before the refusal propagates; the human surface still prints nothing.

   This is a DELIBERATE deviation from `after_generation`'s contract, which says
   the caller "must print NO output" on a forced refusal. Named rather than
   quiet: that rule exists so a CPU result is never read as a GPU success, and a
   document carrying `"backend": {"fell_back": true}` beside exit 14 cannot be
   read that way, while a human-formatted success blob can. The protective half
   is kept; the half that blinded `--json` consumers is not.

2. The caller's `if let Some(note)` arm was UNREACHABLE. `announced` was
   `Some("gpu")` exactly when `forced` was true, and `after_generation`'s
   corrective-line branch requires `forced == false` — so the Some arm could
   never execute. A branch with no reachable caller, which is precisely the
   defect this PR fixes, reproduced one layer down while fixing it.

   `reconcile_accelerator` now returns `Result<()>`, wires the FORCED half only,
   and says in its doc comment why the default-selection half is not wired:
   nothing calls `registry::announce`, so there is no recorded announcement for
   a default selection, and manufacturing one would assert a choice this process
   never made. That is REG-8, not this PR.

Verification: cargo test -p apr-cli --lib = 7289 passed, 0 failed, 12 ignored;
cargo fmt --all --check = 0; cargo clippy -p apr-cli --lib -D warnings = 0.

Refs #3602
Pmat-Ticket: PMAT-3602

* PMAT-3346 (adoption): a bias vector is as wide as the projection it biases

Found by the AD-04 quorum on #3350 (lane 1, gemini-3.1-pro-high, cited
model_arithmetic.rs:144): the projection term used q_out for a gated
family's q matrix (2*n_h*d_k — attn_q emits the output gate, MEASURED in
Qwen3.5-0.8B) while the bias term still used q_dim. A bias narrower than its
projection is not a model. One token: q_dim -> q_out in the bias sum.

Why a delta-0 measurement did not catch it: no shipped family exercises the
case. Qwen3.5 has no attention bias; Qwen2.5 has biases but is not gated, so
q_out == q_dim there. The test that pinned 12 (a q_dim bias under a q_out
matrix) now asserts 16 and says why, and a second test holds the other
polarity — a non-gated family with biases is unchanged at 12.

115 model_arithmetic + model_family tests pass; oracle 218; clippy clean.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(run): quorum round 2 — a --json --benchmark leak, a comment claiming coverage it lacks, and a stray artifact I swept in

Three findings, two lanes, all correct.

1. LANE 1 — `--json --benchmark` leaked a human success blob on a refused run.
   The refusal path carried its OWN copy of "is this a machine surface"
   (`stream || output_format == "json"`) and dropped `!benchmark`. So that
   combination entered the branch, matched neither machine arm inside
   `print_run_output`, and fell through to the human benchmark rendering — for a
   run being refused. Two spellings of one condition drifting apart in the gap
   between them, which is the defect this PR's own dispatch comment warns about.

   One spelling now: `emits_machine_output(stream, output_format, benchmark)`.
   `the_machine_output_predicate_matches_print_run_output` pins it to the arms
   it describes over every flag combination; planting the pre-fix spelling turns
   it and `json_plus_benchmark_is_not_a_machine_surface` RED, verified.

2. LANE 1 — the classifier comment claimed `--gpu-layers all|n` was among the
   flags handled here. It is not: `Commands::Run` carries `gpu` and `no_gpu` and
   nothing else; `--gpu-layers` belongs to `apr serve`. So
   `layers_want_accelerator: false` is CORRECT and the comment was the defect —
   a comment asserting coverage the code does not have. Both comments now say
   the input is genuinely absent for this surface rather than stubbed.

3. LANE 3 — `trace-1789944278.json` (232 lines) was committed to the repo root.
   A runtime byproduct of `test_print_chrome_trace_creates_file`, which writes
   `trace-<epoch>.json` into the CWD when given no `--trace-output`. I swept it
   in with `git add -A` immediately after a quorum round — the exact thing my
   own notes say never to do there. Removed, and this commit stages files by
   name. The test writing into the repo root is a separate defect and is not
   fixed here.

Round 2 was 2 FAIL / 1 NO-VERDICT. Lane 2 has now returned NO-VERDICT twice with
`envelope_status: SUCCESS`, `transport_status: SUCCESS` and non-empty
`raw_bytes` — it answers, and the harness cannot extract a verdict from what it
returns, so width 3 has been delivering two votes.

Verification: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored;
cargo fmt --all --check = 0.

Refs #3602
Pmat-Ticket: PMAT-3602

* roadmap: transcribe #3604 verbatim — round 2's pro lane was right

Criterion (5) had grown an implementation-status clause from PR #3634's
report ('the --json field lands with #3606 StageTimings…') that issue #3604
does not state, and the out-of-scope list was a paraphrase. PMAT-3637 (1)
says every criterion present, none added: the notes now carry the issue's
done_when 1-6, admission rule and out-of-scope list as written.

Refs #3604, #3634
ont-delta: none — roadmap entries only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* audit: quorum-PMAT-3637 — 3/3 PASS, round 3 on the ruled trio

pro / 3.7-flash / 3.6-flash, three conversations, no silences. Rounds 0-2
each said something true about this PR: round 0 judged the wrong ticket;
round 1 found a receipt outside scope and an invented criterion; round 2
found an implementation-status clause in a 'none added' transcription.
None is committed — each judged a diff that no longer exists.

Refs #3604, #3634
ont-delta: none — audit artifact only

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3346: quorum verdict 3/3 on e060c7e2c (AD-04)

Round 0 was 1 FAIL / 1 no-verdict / 1 PASS and the FAIL was real (the bias
width, fixed in e060c7e2c). Round 1 on the fixed head: 3/3 PASS,
gemini-3.1-pro-high / pro-low / 3.6-flash-high, each measured, no dissent.

Refs #3346, #3350

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(audits): AD-04 quorum receipt for PMAT-3577 — AGREED 3/3

Three independent agy lanes on distinct models, none in the author's family
(author measured as Opus 5 / claude):

  gemini-3.1-pro-high  = PASS
  gemini-3.7-flash-high = PASS
  gemini-3.6-flash-high = PASS

Lane models set per-invocation via PAIML_IMPLEMENT_CONFIG rather than by editing
the shared config, which three sessions were launching against concurrently.

THE MERGE DID NOT CHANGE WHAT WAS JUDGED. The lanes ran against 7f82e2cd3; the
PR head is 4fbd2ec31, a merge of main. The trees differ by four files, but all
four arrived FROM main, and the PR's own contribution relative to main is
byte-identical across the merge:

  files, pre-merge  : 49        files, post-merge : 49
  only in post      : (none)    only in pre       : (none)
  sha256(diff vs merge-base), pre-merge  : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35
  sha256(diff vs merge-base), post-merge : bf000a5ea9f94ce5c4b5e7d470f38e84bd6b6dac9760b970c1d5f89945c64d35

So this is the same head for review purposes and does not need a new round. A
naive `git diff origin/main..HEAD | sha256sum` from my local head disagrees only
because that head carries this receipt commit, which the lanes never saw —
comparing the receipt's parent is what makes the two comparable.

Refs #3600
Pmat-Ticket: PMAT-3577

* fix(tests): the shape-count ratchet was red — #3605 adds a 6th shape (quorum round 1)

`the_tracked_repo_graph_is_fresh` asserts `shapes_n == 5` and names the five.
`refusal-receipt-v1` makes six, so this PR shipped a CI red that a quorum lane
found and I did not.

MEASURED, not guessed: `pv extract contracts --check` on this branch reports
`shapes_n: 6, triples: 15620`.

The hardcoded count stays hardcoded. A shape added without anyone noticing is
exactly what this assertion exists to prevent, so adding one is SUPPOSED to turn
it red and make you name the new shape. Noted in the comment that the counter is
shared across branches — a sibling PR adding a shape (#3600's parity-receipt-v2)
will need it raised again at merge, which is the ratchet working rather than a
conflict to route around.

Refs #3605
Pmat-Ticket: PMAT-3605

* chore(audits): AD-04 quorum receipt for PMAT-3602 — AGREED 3/3

gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS.
Author measured as Opus 5 (claude); no lane in the author's family. Lane models
set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared
config that other sessions were launching against.

This trio was chosen on measured odds after flash-class lanes returned
NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my
earlier runs). All 15 lanes voted in this batch.

Refs #3638
Pmat-Ticket: PMAT-3602

* chore(audits): AD-04 quorum receipt for PMAT-3605 — AGREED 3/3

gemini-3.1-pro-high = PASS, gemini-3.1-pro-low = PASS, gemini-3.6-flash-high = PASS.
Author measured as Opus 5 (claude); no lane in the author's family. Lane models
set per-invocation via PAIML_IMPLEMENT_CONFIG, never by editing the shared
config that other sessions were launching against.

This trio was chosen on measured odds after flash-class lanes returned
NO-VERDICT intermittently (3.7-flash 3/7, 3.6-flash 1/7, pro-high 0/7 across my
earlier runs). All 15 lanes voted in this batch.

Refs #3613
Pmat-Ticket: PMAT-3605

* PMAT-3351 (adoption): re-id the fragment — PMAT-3347 is #3348's row on main

Both #3348 (merged; bound the qwen35-e2e equations, closed #3347) and this PR
(the L2 column reads a link, not an index) address issue #3347, and both
claimed roadmap id PMAT-3347. Two PRs cannot share a row: main's PMAT-3347
stays as #3348's, this PR's fragment becomes PMAT-3351 (github_issue 3347,
same title, same notes), and the aggregate is rebuilt from main's roadmap.yaml
plus this branch's fragments — 935 + 1 = 936, nothing re-serialised.

Refs #3347, #3348

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3351: notes transcribe issue #3347's Ask verbatim (bullets 2 and 3; bullet 1 was #3348)

The quorum lanes read pmat work status, i.e. this fragment's notes, not the
GitHub issue. Empty notes would have them judge against a title. The issue
has no done_when section, so its Ask and its two defect statements are
transcribed as written, with the out-of-scope bullet named and no count
offered as a criterion.

Refs #3347

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3338: register the second ticket this PR closes, so the quorum judges the truncate fix as in scope

Quorum round 0 on a4b76397f was 2 FAIL / 1 PASS, both FAILs on one point: the
branch's first commit fixes #3338 (--table panicked on a byte-index cut inside
a multi-byte char) and PMAT-3351 does not ask for it. Both lanes called the
in-scope work correct. The fix is a prerequisite — the L2 column is exercised
through --table on the real corpus, which panicked before it — and #3338 is
an OPEN issue this PR genuinely closes. So it gets its row, transcribed from
the issue, and round 1 runs with --ticket PMAT-3351,PMAT-3338.

Refs #3338, #3347

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3604: the receipt's temp file is private to its writer — a shared name made the atomicity claim false

AD-04 quorum round 0 on #3634, lane 1 (gemini-3.1-pro-high), cited
f2_receipt.rs:180: write_receipt used one shared <sha>.json.tmp, so two apr
runs validating the same model at once could have writer B truncate the file
writer A was about to rename, and A rename B's partial into place. The reader
would classify it Unreadable and validate — the safe direction, done_when 4 —
but the docstring said 'never a truncated one', and that was false.

The temp name now carries the pid and a per-process counter; a failed rename
removes its own temp. A racing test spawns two writers on one path forty
times and parses the survivor each time: it is always one writer's WHOLE
receipt. 14 tests.

Lane 1's other two findings, for the record: --revalidate is done_when 2
verbatim (the fragment carries it), not scope creep; and F2Outcome's fields
are the values the stderr line already prints and done_when 5's data path
for #3606's JSON, not dead code. Two lanes passed the same diff.

Refs #3604

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3604: quorum verdict 3/3 on dd546592b (AD-04)

Round 0 was 2 PASS / 1 FAIL; the FAIL was the shared temp-file race, fixed in
dd546592b. Round 1 on the fixed head: 3/3 PASS, gemini-3.1-pro-high /
pro-low / 3.6-flash-high, each measured, no dissent.

The PMAT-3604 roadmap row was on this checkout UNCOMMITTED for the resolver
(pmat work status reads the checkout); it lands via #3637. The lanes judged
the committed diff origin/main...HEAD.

Refs #3604, #3637

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* PMAT-3351 + PMAT-3338: quorum verdict 3/3 on d915b8f16 (AD-04)

Round 0 (--ticket PMAT-3351 alone) was 2 FAIL / 1 PASS, both FAILs on the
#3338 truncate fix being out of scope; both called the in-scope work correct.
Round 1 with both tickets registered and named: 3/3 PASS, gemini-3.1-pro-high
/ pro-low / 3.6-flash-high, each measured, no dissent.

Refs #3347, #3338

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(contracts): the v1 contract had no equations, and its own denominator guard was never wired

Two reds on this PR, both mine, and the second is the better finding.

1. `contract_data_integrity` is a SHRINK-ONLY ratchet at 444 and this branch made
   it 445. The 445th is `parity-receipt-v1: no equations`. Measured which one
   rather than guessed — `parity-receipt-v2` is not in the list.

   `kind: pattern` correctly declares no SHAPE (a shape over a class nothing
   instantiates passes vacuously, which is #3610's whole lesson), but equations
   are not shapes, and this contract does assert things: PRC1-INV-001 and
   PRC1-INV-002 already said them. They are now written as the two equations the
   integrity check reads — `no_instances` and `refused_by_name` — rather than as
   new claims invented to satisfy a counter.

2. `check_guards_are_wired.sh`: `NEW: parity_receipt_denominator.sh`. I ADDED A
   GUARD IN THIS PR AND WIRED IT INTO NOTHING — the second, independent reader of
   the receipt corpus, shipped where no workflow names it.

   Minutes before finding this I wrote into that same contract's equations, as a
   precondition: "both readers are wired into CI — a refusal nothing runs is not
   a refusal." My own precondition, violated by my own PR, in the same file.
   `check_guards_are_wired.sh` caught what I had just finished writing down.

   Now named in `ci.yml` beside the other cargo-free, model-free case tables. Its
   4-case self-test passes: a legacy record refused by name, a receipt added
   without bumping the denominator disagreeing, bumping it making them agree.

NOTE — THIS TOUCHES `.github/workflows/ci.yml`, which CLAUDE.md lists as a
check-in item. Additive only: one step in an existing guard block, no matrix,
trigger or gate logic changed. Wiring an unwired guard is the minimum the failing
check asks for, and leaving it unwired to avoid the file would be keeping a
refusal that never runs.

VERIFICATION
  cargo test -p aprender-contracts --test validate_contracts contract_data_integrity   1 passed
  bash scripts/parity_receipt_denominator.sh --self-test                               rc 0, 4/4
  bash scripts/check_guards_are_wired.sh                                               PASS (4 -> 3)
  pv validate contracts/parity-receipt-v1.yaml                                         0 errors, valid

DIFF MOVEMENT, for the receipt: the judged diff DID change — an equations block
and one CI step. No behaviour the lanes reviewed was altered; the extractor, the
shapes and the denominator script are byte-identical. AD-04 re-review is the
inspector's call.

Refs #3577
Pmat-Ticket: PMAT-3577

* PMAT-3401 (adoption): the three derivatives 24 new contracts oblige — census, graph, README — regenerated together

This PR adds 24 contracts and regenerated none of the tracked artifacts
derived from the corpus. On #3581 that omission surfaced one per CI round,
each masked by the one before it. All three here, at once, with a pv built
from this tree under a pinned target dir:

  contracts/census.json   1800 -> 1824  (+24, the contracts added)
  contracts/contracts.nt  15,600 -> 15,696 triples  (+96 = 24 x 4; GREW —
                          a drop is the tell for a malformed input or a stale
                          binary; binding.yaml still parses, 156 entries)
  README CONTRACT_COUNT   2 blocks -> 1824 via make readme-sync

The merge took main's generated README blocks over the branch's hand-typed
1866, then readme-sync wrote the measured 1824; the branch's number described
a tree that never existed on main.

Verified: all 24 contracts pv-validate under the pinned binary (control:
main's crux-A-01 valid under the same one); lint_passes_on_real_contracts
green, so the sigma prose ratchet holds; 1666 engine tests; ont4b shapes gate
11/11; test-binding ratchet; readme-sync-check; FALSIFY-README-002; roadmap
additive added=26 deleted=0, aggregate idempotent, 26 fragments present.

Refs #3401

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(run): extract reconcile_and_emit — run() hit cognitive 27 against a 25 ceiling

`check_complexity_ratchet.sh` went RED with:

  RED  NEW  crates/apr-cli/src/commands/run_entry.rs::run  cyclomatic 21 cognitive 27
            (over a threshold, absent from the comparand)

The reconcile-then-emit block I added in this PR is what took `run` over. Moved
into `reconcile_and_emit`, which also gives the ordering decision — machine
surfaces emit on a refusal, the human surface does not — a place to be documented
that is not the middle of a 30-argument entry point.

Now: PASS (D2): e6f77c98c vs 19750b384 measured by pmat 3.41.1 — none new, none grown.

NOTE FOR THE NEXT PERSON: I first re-ran the ratchet against my WORKING TREE and
saw the identical cognitive 27, and briefly concluded the extraction had not
helped. It had. The script measures two REVISIONS — it prints
`merge HEAD <sha>` — so an uncommitted fix is invisible to it. Commit, then
measure.

Tests: cargo test -p apr-cli --lib = 7291 passed, 0 failed, 12 ignored;
fmt --check = 0; clippy -D warnings clean.

Refs #3602
Pmat-Ticket: PMAT-3602

* PMAT-3496: a registration row for the registration PR — the author's earlier receipt named this ticket but no row existed on the branch

* PMAT-3496: clause (1) transcribes issue #3495's title verbatim (132 Kani harnesses; Verus pilot), not a paraphrase

* fix(guard): a step NAME read as a bare guard_tree.sh dispatch hid four dark guards; the #3305 guard runs nowhere and dies on intel's python 3.10 (#3644)

Refs #3644 #3646 #3305 #3626

THE FINDING, proven by mutation before it was argued. check_guards_are_wired.sh
counts a guard wired-by-dispatch when a workflow runs guard_tree.sh in a mode
whose --dry-run RUN set contains it. Its invocation regex accepts `guard_tree.sh`
followed by `'` (for a quoted run: string). ci.yml has two step NAMES,
`"guard_tree.sh's own case table …"` and `"guard_tree.sh's parallel dispatcher
…"`; the apostrophe matched, neither line says --no-cargo, so each was read as a
bare dispatch in mode "all" -- 55 cargo-classified guards counted wired that
`guard-tree` (--no-cargo) never runs and `guard-cargo` never names. Rewriting
those two names (nothing else) turns the meta-guard RED: 3 -> 7 unwired --
check_model_ladder.sh, check_pathonly_devdeps_unused_in_src.sh,
check_pr_review_counts.sh, check_pr_review_receipt.sh. Three of the four are
cargo-classified by a COMMENT that says the build tool's name; none of those
three invokes it.

check_pathonly_devdeps_unused_in_src.sh (#3305: src/ may not use a dev-dep that
publishing deletes; clean-room red 8/8 across two releases) has therefore run
nowhere since it was written on 2026-09-15. And it could not have run on the
intel hosts: `import os, re, sys, tomllib` at module level, python 3.10.12
there with neither tomllib nor tomli (measured on mac-server tonight; the
sovereign-ci container has no python3 at all). Reproduced with python3.10 -S:
the traceback is swallowed by `out=$(scan … || true)` and the case table prints
"FAIL: missed hit_use" ×4 -- the regex reported broken by an interpreter that
never ran it. Same class as #3626's guard on intel-clean-room-6.

WHAT CHANGES

scripts/check_guards_are_wired.sh
  - `_not_a_name_line` drops `name:` lines before the invocation match, in
    dispatcher_wired() and the per-guard scan. A step name is documentation
    whatever it contains.
  - rows 8-10: the exact ci.yml fixture (`- name: "guard_tree.sh's own case
    table …"`, no run: line) must leave the dispatched guard unwired; a QUOTED
    run: line still dispatches (the control the `'` exists for); a step named
    after a guard wires nothing. Mutant (drop removed): rows 8 and 10 RED.
  - the ledger's ratchet is now set-aperture, owned by this file.

scripts/lib_baseline_ratchet.sh, scripts/check_baseline_ratchets.sh
  - set-aperture gains a NAME-entry admission for sets of FILES: an added entry
    with no `:` is admitted iff the comparand carries that file (at the entry's
    path, or beside the owning guard -- the ledger names siblings by basename
    and guard_tree.sh reads it so) AND the owning guard changed in the diff.
    A file this branch created is refused (PERF-028's shape). Six rows: predates
    -> admitted; branch WROTE it -> refused; no guard edit -> refused; escaping
    path -> refused; beside the OWNER -> admitted; beside a different owner ->
    refused. Lib mutants (always admit / branch removed / sibling resolution
    removed) each turn rows RED; a separate absolute-path check survived its
    mutant because `git cat-file -e` already refuses those, so it is not there.
  - classify: unwired_guards_baseline.txt -> set-aperture, owner named.

scripts/unwired_guards_baseline.txt  3 -> 6, as APERTURE REVEALS with a reason
  per line (each may only leave): check_model_ladder.sh (T-2 release gate, the
  multiplatform_dogfood class); check_pr_review_counts.sh (RED on main today,
  "6 row(s) disagree" -- #3646); check_pr_review_receipt.sh (takes a receipt
  path; nothing passes one -- #3646). check_pathonly_devdeps_unused_in_src.sh
  is NOT ledgered: see below.

scripts/check_pathonly_devdeps_unused_in_src.sh
  - readers tomllib -> tomli -> ENV rc=2 naming python version and $RUNNER_NAME.
    No purpose-built manifest reader: a second TOML implementation over 78
    manifests, where a wrong "has source" is a silent PASS on exactly the
    #3305 class, is not a fallback worth shipping. UNMEASURED-with-a-name is
    the honest state on a runner with no TOML library.
  - the scanner ends with `SCAN-DONE manifests=N reader=…`; the shell REQUIRES
    it. Output without it (no reader, /bin/false, no python3) is ENV rc=2 --
    never rc=1, never "no violations". The selftest and the tree scan
    propagate 2 instead of swallowing it.
  - five death rows: both readers absent (sys.modules nulled so the imports
    fail for real); interpreter exits 1 silently; no interpreter (127); the
    working interpreter as control; and THE WHOLE GUARD under the intel shape
    (nested, recursion-guarded) -- the call site is where `|| true` hid it.
    Mutants: trailer not required (3 rows RED); trailer printed on the
    no-reader path (1 RED); `|| true` restored (the whole-guard row RED).
  - three comments and one FAIL string reworded so the file no longer says
    the build tool's name with a space after it: it invokes none, and that
    substring is what CARGO_RE classifies on. guard_tree.sh --dry-run
    --no-cargo now lists it as `run:`; bare run on this tree rc=0 in 3 s.

Passes: check_guards_are_wired --self-test 10/10; check_baseline_ratchets
--self-test 61 rows; pathonly --selftest under py…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stale APR-RELEASE-001 §6: >=2 release trains with no activity

Projects

None yet

Development

Successfully merging this pull request may close these issues.

QE2E-INV-001 asserts Qwen3.5-9B ∈ [9.0B, 9.2B] but the repo's own 9b descriptor sums to 8.209B — the descriptor declares no DeltaNet tensors

1 participant