Skip to content

Inference-grade continuous outer choice (GSS) for NNBEGM - #407

Merged
hmgaudecker merged 1992 commits into
mainfrom
feat/continuous-outer
Aug 31, 2026
Merged

hmgaudecker merged 1992 commits into
mainfrom
feat/continuous-outer

Conversation

@hmgaudecker

@hmgaudecker hmgaudecker commented Jul 20, 2026 •

Copy link
Copy Markdown
Member

Implements continuous outer choice for NNBEGM. In economic terms, this changes how one continuous adjustment margin is solved; it does not change the model's preferences, constraints, or transition equations. The finite-grid method remains available.

Economic example

Consider the Mahler–Yum model. A household enters the period with wealth and last period's effort habit, say 0.40. It can:

  1. keep effort at 0.40 without paying an adjustment cost; or
  2. pay an observed adjustment cost, choose a new effort level anywhere between 0 and 1, and then choose consumption and saving conditional on that effort.

With a 17-point effort grid, the adjuster can choose only 0, 0.0625, 0.125, and so on. If its economic optimum is 0.43, the finite solver must publish a nearby grid point such as 0.4375. The continuous solver instead evaluates exact conditional consumption-saving problems at selected effort nodes and adaptively refines promising intervals between them.

There are two distinct simulation cases:

  • With a deterministic keeper-versus-adjuster comparison, simulation interpolates the conditional policies at each subject's actual state and reruns the safeguarded outer search. Separately solved values can be replayed with the exact policies returned by solve(return_simulation_policy=True); supplying NNBEGM values without their matching policy fails closed.
  • With Mahler–Yum's observed fixed adjustment cost, solution integrates the uniform cost analytically. A simulated subject still needs a realized draw to determine whether it keeps or adjusts. Until that draw-and-replay channel exists, simulate() now fails before an automatic solve with a dedicated unsupported-operation error; it cannot silently substitute the ordinary-grid decision.

The analytic integration supplies the expected value and exact adjustment probability at each solved state. It does not replace the subject-level cost realization needed for path-dependent simulation outcomes.

Benefits and costs

Benefits:

  • Less grid-snapping error in the solved outer action.
  • Smoother policy responses for comparative statics, moments, and—in smooth regions—derivatives.
  • The inner consumption-saving problem retains NB-EGM's exact treatment of declared kinks, jumps, and hard constraints.
  • A supported uniform observed adjustment cost is integrated analytically during solution, avoiding another solve-state dimension and yielding the adjustment probability in closed form.
  • Numerical diagnostics expose approximation error, bound solutions, close competitors, and fallbacks.

Costs and limitations:

  • Continuous search requires more conditional solves, compilation, memory, and runtime than a fixed outer grid. Large models need batching or streaming.
  • Results depend on mesh and tolerance settings; difficult value surfaces may exhaust the node budget.
  • fail_closed=True covers mesh convergence specifically: an exhausted node or round budget with marked intervals outstanding raises. Bound optima, close branch or runner-up margins, policy fallbacks, and all-invalid cells remain diagnostics; a broader inference release gate is the caller's responsibility.
  • The analytic adjustment-cost fold currently supports only the declared uniform observed-cost specification and validates every evaluated scale as finite and nonnegative.
  • Fixed-cost lifecycle simulation still needs cost draws and contingent keeper/adjuster replay; simulate() explicitly refuses this configuration in the meantime.
  • The continuously replayed outer proposal interpolates the inner policy between solved outer nodes. Canonical rescoring must agree one-sided with the policy surrogate's winning value within the configured mesh absolute/relative tolerances; otherwise simulation emits the safe baseline and records nested_policy_fallback.
  • Implicit derivatives are available only for resolved, smooth, interior optima.
  • FiniteOuterGrid remains the simpler choice for quick runs, historical comparisons, and debugging.

User-facing API

Import the solver and its strategies from lcm.solvers.

  • NNBEGM(..., outer_search=...) requires an explicit outer-search strategy.
  • FiniteOuterGrid(...) preserves finite candidate-grid behavior.
  • AdaptiveOuterMesh(...) performs the continuous-outer approximation.
  • DeterministicOuterMaximum() keeps the hard keeper/adjuster maximum and gives the keeper exact ties.
  • UniformObservedFixedCost(...) analytically integrates a uniform fixed adjustment cost during solution. It requires AdaptiveOuterMesh, publishes the analytic adjustment probability, and rejects negative, NaN, or infinite evaluated scales before solving.

The branch-aggregation slot is a closed set, not a polymorphic extension point: NNBEGM rejects any OuterBranchAggregator subclass other than the two implemented folds. Inert outer-search options and the unwired LegacyGoldenSection configuration have been removed.

Numerical contract

The continuous path materializes the exact conditional NB-EGM solution at every outer mesh node, constructs a shared validated cubic-Hermite surrogate, identifies every node-local candidate bracket, and refines each bracket with safeguarded golden-section maximization. It does not assume global unimodality.

Exact mesh nodes remain in the candidate set, so refinement cannot lose the finite-grid winner. Candidate selection retains deterministic keeper/adjuster, tie, NaN, and fallback semantics. The numerical audit covers exponent extremes, subnormals, blocking, and both precisions. Only the framed-arithmetic helpers used by the live NB-EGM path remain in production; the parked 1,458-line alternative envelope implementation and its private tests were removed.

The solver publishes SolverDiagnostics, including normalized interpolation error, final bracket width, mesh size, bound masks, branch and runner-up margins, fallback masks, unresolved cells, and analytic adjustment probabilities where applicable.

For deterministic hard-maximum models in the supported reader scope, the nested simulation payload carries the outer candidate bank and conditional inner policies. Simulation reconstructs candidate actions at the subject's actual state, reruns the safeguarded outer search, and scores emitted pairs under canonical Q. A proposal is accepted only when feasible, weakly better than the admissible baseline, and not materially below the policy surrogate's winning value under the configured absolute/relative value band. Otherwise simulation emits the canonically rescored baseline and records nested_policy_fallback.

The baseline it has to beat is the best action the subject could actually be given, not one privileged candidate. Where the solved policy cannot be reproduced at a subject's own state, the fallback ranks the raw grid pair, the no-adjustment keeper, and every published outer node under the same canonical Q — flow payoff plus discounted continuation at that subject's state — and emits the highest survivor. Infeasible and non-finite branches are excluded before the comparison, not after it; each branch carries its own conditional inner policy, so an outer node is never paired with another node's consumption rule; ties are broken deterministically, with the grid pair taking an exact tie and the keeper preceding ascending node order within the bank. If nothing survives, simulation raises before emitting an action rather than publishing an infeasible one. The values the solve published for those branches are not consulted at this point: they ranked branches at the solved state, which is not the state being simulated.

One consequence is worth stating plainly, because it changes simulated behaviour rather than only internals. A maximum over candidates is weakly higher than the grid pair alone, so the interpolated proposal is weakly harder to accept and weakly more subjects are served by the fallback. The direction follows from the acceptance inequality; the size of the effect on any particular model's moments is not characterised here.

Both outer searches recover the outer action by inverting a declared affine map whose
coefficient is certified from the map's structure, and admit the result through one shared
tolerance-free predicate: the forward image must be finite and inside the published inverse
domain, and where the target is a declared endpoint the image must reproduce it bit-for-bit.
There is no residual tolerance and no slope probe on either route, so solve and simulation
cannot gate on different predicates. An outer grid reaching outside the outer state's
declared domain is refused at build time. A candidate the predicate refuses is replaced by the
best admissible branch of the published bank — another outer node, the keeper, or the
grid pair — rather than by a stock displaced from the endpoint.

The optional implicit-derivative primitive uses a mesh-plus-golden-section primal and an implicit-function custom JVP only for resolved smooth interior optima. Boundary solutions, flat curvature, competing basins, and nonstationary/kink optima are marked unresolved rather than assigned a misleading derivative.

Mahler–Yum integration

create_mahler_yum_model(implementation="paper") uses:

  • continuous effort as the NNBEGM outer action;
  • analytic integration of the observed fixed adjustment cost during solution;
  • consumption as the inner Euler action;
  • retirement feasibility in flow utility; and
  • a declared minimum-resource floor.

Two limitations remain explicit:

  1. Paper mode does not yet draw the analytically integrated adjustment cost during simulation and replay the contingent keeper/adjuster policy. simulate() raises a dedicated unsupported-operation error for this configuration until that channel exists.
  2. The authors' Fortran uses c = max(cash_on_hand - saving, min_consumption). Once its floor binds, every additional dollar saved increases the implied government transfer dollar-for-dollar; with no asset test, the household may preserve all its own resources while government finances the entire consumption floor. The declared pylcm model instead lets consumption and saving divide max(cash_on_hand, min_consumption). The implementations therefore differ only in below-floor states, and the difference is documented rather than hidden.

The incomplete legacy_fortran compatibility layer has been deleted. The factory now offers only the paper-equation model and the brute-force oracle; it does not expose switches that cannot truthfully describe the constructed model.

Scope

This PR contains the continuous-outer solver, deterministic nested-policy simulation reader, diagnostics, and paper-equation Mahler–Yum model. Fixed-cost draw-and-replay remains unimplemented but is now a fail-closed unsupported simulation operation; materially inconsistent interpolated inner policies are controlled by the configured value band and explicit safe fallback. Moment construction, parameter transforms, standard errors, estimation release gates, and published-number decomposition remain outside this PR.

Verification

Head: a5c6fe4fc1c70354cf16efd341c419a8e8975fb7.

  • Fallback selection, tests/simulation/test_outer_endpoint_admission.py: 37 tests, 0
    failures and 0 errors
    at --precision=64 and at --precision=32. The module is a
    discriminating instrument, not merely a passing one: run against the previous selection
    rule with the same tests, 5 of the 37 fail at both precisions.
  • Resource cost of ranking every candidate rather than one, measured on the adaptive
    workload with the two rules alternated across three blocks per precision so that machine
    drift cancels within a block: warm simulate +0.5% (f64) and +0.1% (f32),
    best-of-7; cold solve+simulate within [-0.6%, +1.5%]; peak resident memory within
    [-0.7%, +2.2%]. All inside the 5% warm and 10% compile/memory review thresholds.
    Alternating mesh shapes reuse compiled programs: the compilation-cache entry count is
    flat across eight repetitions on both rules.
  • prek run --all-files exit 0: 34 hooks run, 34 passed, 0 failed, 0 skipped, including
    Ruff, formatting, codespell, and ty 0.0.75. The 35th declared hook, pixi-lock-check,
    runs at the pre-push stage.

Not claimed here: the full suite at this head. CI covers it — the cpu context is green
only if all of the fp64 leg, the fp32 leg, and three sharded slow-solution jobs at both
precisions succeed, alongside macOS and Windows platform legs and the exact-kernel
capability matrix — and its result on this head is what should be read, not a number
transcribed into this description.

API change for downstream branches

NNBEGM no longer accepts branch_aggregator; it now carries exactly
{inner, outer_search}. The adjustment cost is declared on the margin instead:

OuterContinuousMargin(
    ...,
    adjustment_cost=UniformObservedFixedCost(...),
)

Leaving adjustment_cost unset means adjusting is free at the margin, i.e. the
deterministic maximum. UniformObservedFixedCost still requires
outer_search=AdaptiveOuterMesh(...), now reported against the regime that
declared it rather than against the solver.

Nothing on main is affected — branch_aggregator was introduced in this PR and
never existed on main. Branches stacked on this one inherit the migration with
their next merge; any NNBEGM(branch_aggregator=...) written on such a branch
before taking that merge becomes a TypeError.

@hmgaudecker

Copy link
Copy Markdown
Member Author

Status note: the base branch moved while this PR was being assembled — feat/nested-nbegm-ez merged the dcegm solver-boundary refactor (round-15 repairs) at 03:16, which splits solvers.py into per-solver modules (nbegm.py, nnbegm.py, …). This PR's 18 commits modify the pre-split solvers.py, so the PR currently shows conflicts.

The port is mapped and tractable: ~907 diff-lines of solver delta to re-route into the new modules (everything else — the new _lcm/egm/ modules, lcm_examples paper mode, tests — should merge clean). Plan: a single git merge conflict session rather than a rebase, then the frozen-baseline neutrality test, the NBEGM/NNBEGM batteries, the Mahler merge gates, and the full CPU suite. The base merge should also fix the two pre-existing test_n_nbegm failures documented in the PR body — they are expected to flip green after the port.

🤖 Generated with Claude Code

@hmgaudecker hmgaudecker changed the title Inference-grade continuous outer choice for NNBEGM (plan PRs 0-8) Inference-grade continuous outer choice for NNBEGM (plan PRs 0-8, 12, 13) Jul 20, 2026
@github-actions

github-actions Bot commented Jul 20, 2026 •

Copy link
Copy Markdown

Benchmark comparison (main → HEAD)

Comparing 9987966f (main) → a288bff8 (HEAD)

Benchmark Statistic before after Ratio Alert
aca-baseline execution time 9.987 s 9.999 s 1.00
peak GPU mem 601 MB 601 MB 1.00
compilation time 423.51 s 417.83 s 0.99
peak CPU mem 10.59 GB 10.01 GB 0.95
aca-baseline-debug execution time 34.756 s 34.433 s 0.99
peak GPU mem 603 MB 602 MB 1.00
compilation time 470.67 s 456.57 s 0.97
peak CPU mem 6.43 GB 6.46 GB 1.00
Mahler-Yum execution time 2.992 s 2.985 s 1.00
peak GPU mem 685 MB 685 MB 1.00
compilation time 22.66 s 22.30 s 0.98
peak CPU mem 1.50 GB 1.49 GB 1.00
Precautionary Savings - Solve execution time 25.4 ms 26.6 ms 1.05
peak GPU mem 34 MB 34 MB 1.00
compilation time 2.42 s 2.45 s 1.01
peak CPU mem 1.07 GB 1.07 GB 1.00
Precautionary Savings - Simulate execution time 62.8 ms 65.0 ms 1.04
peak GPU mem 328 MB 327 MB 1.00
compilation time 4.95 s 4.98 s 1.01
peak CPU mem 1.23 GB 1.24 GB 1.00
Precautionary Savings - Solve & Simulate execution time 98.3 ms 100.6 ms 1.02
peak GPU mem 1.68 GB 1.68 GB 1.00
compilation time 7.16 s 7.37 s 1.03
peak CPU mem 1.26 GB 1.26 GB 1.00
Precautionary Savings - Solve & Simulate (irreg) execution time 234.9 ms 237.4 ms 1.01
peak GPU mem 2.20 GB 2.20 GB 1.00
compilation time 7.32 s 7.32 s 1.00
peak CPU mem 1.31 GB 1.32 GB 1.00
IskhakovEtAl2017DCEGMSimulate execution time 219.7 ms 237.6 ms 1.08
compilation time 7.66 s 7.58 s 0.99
peak CPU mem 1.46 GB 1.47 GB 1.01
IskhakovEtAl2017DCEGMSolve execution time 1.900 s 1.896 s 1.00
compilation time 7.25 s 7.11 s 0.98
peak CPU mem 1.37 GB 1.38 GB 1.01
IskhakovEtAl2017Simulate execution time 233.0 ms 228.6 ms 0.98
compilation time 7.56 s 7.58 s 1.00
peak CPU mem 1.22 GB 1.21 GB 1.00
IskhakovEtAl2017Solve execution time 55.4 ms 54.8 ms 0.99
compilation time 2.61 s 2.60 s 1.00
peak CPU mem 1.12 GB 1.12 GB 1.00
IskhakovEtAl2017DCEGMSimulateGpuPeakMem peak GPU mem 440 MB 443 MB 1.01
IskhakovEtAl2017DCEGMSolveGpuPeakMem peak GPU mem 0 MB 0 MB 1.03
IskhakovEtAl2017SimulateGpuPeakMem peak GPU mem 440 MB 442 MB 1.01
IskhakovEtAl2017SolveGpuPeakMem peak GPU mem 67 MB 67 MB 1.00

@hmgaudecker hmgaudecker changed the title Inference-grade continuous outer choice for NNBEGM (plan PRs 0-8, 12, 13) Inference-grade continuous outer choice (GSS) for NNBEGM Jul 20, 2026
Base automatically changed from feat/nested-nbegm-ez to feat/nb-egm July 21, 2026 11:14
@hmgaudecker
hmgaudecker changed the base branch from feat/nb-egm to feat/perceived-stochastic-transitions July 21, 2026 12:10
hmgaudecker added a commit that referenced this pull request Jul 21, 2026
Addresses five of the round-3 full-branch Pro findings on PR #407, the
self-contained ones outside the nested-simulation payload:

- F6 (DC-1, serious): the query.py upper-envelope value-tie test used a
  fixed absolute half-width `_VALUE_TIE_ATOL = 1e-12`, precision-blind at
  large value magnitudes. In float32 near 1e6 (and even float64 near 1e4)
  one ULP dwarfs 1e-12, so two branches at a genuine value tie were split
  and the right-continuous winner was reversed. Replaced with a
  dtype+magnitude-scaled band `_TIE_BAND_ULPS * eps * max(|a|,|b|)` in both
  the dense and blocked paths. Regression in float32 and float64.

- F7 (serious): safeguarded_continuous_argmax capped refinement at the
  top-8 node-local maxima; a lower-ranked basin hiding the tallest off-node
  interpolant peak had that peak silently discarded. Default is now
  max_brackets=None (refine every node-local maximum); an explicit integer
  stays available as a performance knob with the off-node risk documented.
  Regression: a nine-basin surface whose ninth basin holds the global peak.

- F9 (moderate): the optional Lipschitz mesh safeguard chased a
  finite/nonfinite feasibility boundary forever (endpoint_max stays finite
  off the one live endpoint). Conjoined the upper-bound mark with both
  endpoints finite so boundary intervals fall through to the ordinary rules.

- F10 (minor): max_outer_interpolation_error is a dimensionless normalized
  validation ratio, not a value-unit gap; relabeled.

- F11 (minor): the SimulationPolicy seam text claimed a solver-supplied
  reader is the only consumer that looks inside, but the engine's
  simulation loop dispatches on the concrete payload type over a closed
  union; contract text revised to describe the deliberate closed-union
  dispatch.

Verified: 35 outer-refinement + audit-regression + singleton tests green;
the two new regressions green, against the branch src via PYTHONPATH.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Jul 28, 2026
…us-outer (#407)

59 commits, bringing the whole feat/nb-egm line -- including the MSS round-16
exact segment-envelope rewrite and the new shared `double_double` /
`certified_sign` modules -- onto continuous-outer.

Seven conflicts, resolved on what the code MEANS, not on which side is newer:

* `regime_building/diagnostics.py` -- union of the keyword-only params, then
  DROPPED my `continuation_grid_signature`: the merged body derives the
  signature via `continuation_grid_signature_from_schedule(grid_schedule=...)`,
  so my parameter was dead and callers passing it would have been silently
  ignored. Upstream generalized my hook to one source of truth (the schedule).
* `regime_building/processing.py` -- theirs throughout. My `_process_one_function`
  docstring described `_SpecializedEconFunction`, which the merged body no longer
  contains (upstream replaced the marker with `PeriodizedUserFunction`
  normalization, addressing the same age-specialized concern earlier in the
  pipeline). Verified `_SpecializedEconFunction` has no remaining references.
* `egm/nbegm_step.py` -- theirs. `_next_segment_id` is MOVED, not deleted; and my
  `_DEGENERATE_MARGINAL_TOL` is genuinely dead because the auto-merged body
  already replaced all four `marginal <= TOL` sites with upstream's
  `_degenerate_inversion(marginal=, consumption=)` predicate -- a predicate, not
  a magnitude threshold.
* `solution/nbegm.py` -- theirs. `str` -> `StateName`, and my
  `space.state_names[0]` fallback is exactly what upstream's "Resolve the NB-EGM
  liquid axis once instead of assuming state_names[0]" removes.
* `solution/nnbegm.py` -- MINE plus their new guard. Taking theirs would have
  deleted `resolved_outer_search` (used twice) and the UniformObservedFixedCost
  gate, neither of which upstream has; their one line was the SAME guard call
  already in `__post_init__`, whose real change was `solver_name="NNBEGM"` so the
  message stops naming NEGM -- applied to the existing call instead of duplicated.
  Kept my policy-leaf strip (audited round-4 F1: the NNBEGM cross-period
  continuation republishes the policy-FREE bridged collapse) and added their
  `_fail_if_inner_carry_rows_not_grid_aligned`.
* `simulation/simulate.py` -- UNION, and the one that would have broken silently.
  The two sides are orthogonal: mine adds the `NestedEGMSimPolicy` read and the
  nested-fallback flag (the `nested_policy_fallback` field), theirs adds
  canonical-Q scoring for the discrete-branch redecide. Upstream folded the guards
  into `if sim_policy is None or read is None: return`, which would SKIP the
  nested dispatch whenever `egm_policy_read` is unset -- disabling the
  continuous-outer path this branch exists for. Kept my ordering (nested dispatch
  BEFORE the `read` check), my tuple return, and their `score_actions`/`grid_values`.
* `tests/solution/test_envelope_query.py` -- union; disjoint test classes.

Text merged clean but broke semantically in two further places, both fixed here:
upstream made the case-piece declaration API keyword-only, and
`lcm_examples/mahler_yum_2024/paper.py` still called `piecewise_affine("...")` and
`affine_breakpoint("...")` positionally (collection error in three test files). An
AST scan over src/ and tests/ confirms no positional call to a keyword-only
case-piece helper remains. The NBEGM error message advertising the old positional
form was corrected too, since a user following it verbatim would now get a
TypeError.

Two duplicate-parameter/keyword bugs from my own first-pass resolution were caught
by compiling every file -- `ast.parse` does NOT reject a repeated keyword argument,
only `compile` does.

VERIFICATION -- per file across the WHOLE tests/ tree (324 files, 0 UNVERIFIED).
One process over all of tests/ was OOM-killed by cap's 48G at 97% with no summary;
per file bounds the peak and makes a missing summary detectable instead of
mistakable for a pass.

  tests/solution, tests/simulation, tests/egm, tests/regime_building: ALL CLEAN.
  The seven long-standing NEGM/EGM failures are FIXED by this cascade, as the
  stale-nb-egm diagnosis predicted.

38 tests fail in five root-level tests/*.py files. They are NOT caused by this
merge and NOT present on #406 -- measured in fresh detached worktrees with
`_lcm.__file__` asserted:

  #406 @ 68d871f          149 passed, 2 skipped, 0 failed
  pre-merge @ 5fd6cde     38 failed (30 + 8)
  merged                  38 failed (30 + 8), identical

So feat/continuous-outer introduced them and this merge is neutral on them.
Reported separately; not fixed here, because they are a pre-existing branch defect
of their own and not part of a cascade.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Jul 29, 2026
Absorbs upstream #407, which itself carries the #406 perceived-stochastic-
transitions merge (c0699e4). The substantive upstream change for this branch is
the age-normalization refactor: `_resolve_age_specialized_state_grids` became
`normalize_age_specialization` + `grid_schedule`, `resolve_specialized_nodes(x,
age)` became `resolve_periodized_nodes(x, representative_period)`, and
`_grid_identity` / `_node_fingerprint` / `_continuation_grid_signature` are gone.

Conflicts resolved by provenance, not by side-taking:

* `regime_building/processing.py` (8 conflicts) — took upstream's normalization
  pipeline wholesale; kept `regime_to_v_interpolation_info_for_Q`, the collective
  E1 `compute_intermediates` guard, the `get_Q_and_F_collective` branch and the
  four gated-edge fold helpers. The collective branch is now ported to
  `resolve_periodized_nodes(..., representative_period)` so it resolves functions
  and constraints exactly as the singleton branch does.
* `regime_building/Q_and_F.py` — upstream widened
  `QAndFFunction.next_regime_to_V_arr` from `FloatND` to
  `MappingProxyType[RegimeName, FloatND]`. Auto-merge updated the singleton
  builders only; the two collective twins (lines 735, 1285) were left behind and
  are fixed here. `ty` caught this, not the test suite.
* `tests/regime_building/test_collective_phase_role_class_invariant.py` — the
  test pinned the literal `resolve_specialized_nodes(continuation_functions,
  age)`, which upstream renamed. Production was correct; the assertion is
  re-expressed through a derived `_resolver_name()` helper so it tracks the
  resolver rather than its spelling.

Attribution against the c0699e4 baseline (fresh worktree, PYTHONPATH override,
`_lcm.__file__` asserted, same pixi env):

* `tests/solution` — 1212 passed, 3 skipped, 1 xfailed, 0 failed.
* rest of `tests` — 38 failed + 7 errors, an EXACTLY identical set at the
  baseline (`comm` empty in both directions). All 38 are one upstream defect:
  d0d655c added the `nested_policy_fallback` simulation column without
  regenerating the pickled fixtures, so every failure is a one-column DataFrame
  shape mismatch. The 7 errors are the separate AdaptiveOuterMesh OOM.

No new failures are introduced by this merge.

Note for anyone attributing failures on this stack: `beartype_package()` caches
INSTRUMENTED bytecode as `__pycache__/*.opt-beartype<ver>.pyc`, keyed on the
beartype version and not on the conf. A merge can leave a `.pyc` compiled under a
different conf — `is_pep484_tower` off — which is NEWER than its source, so no
mtime check catches it. That produced 120 phantom `tests/solution` failures here
(`int 0 not instance of float` in code byte-identical to the passing baseline).
Clear `src/**/__pycache__` before believing any such result.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01M4Vd9xVvBDhAzoxqGoriPn
hmgaudecker added a commit that referenced this pull request Jul 31, 2026
Propagates the age-specialized/EGM continuation-grid cascade fix into
feat/continuous-outer. Five files conflicted; #400 and #407 have independently
rewritten the query-side envelope certification with no shared symbols beyond
the entry point and `_SegmentLinks`.

Per the resolution recorded in
/tmp/pylcm-fleet-coordination/envelope-architecture-390-400-407.md, #407's
exact-dyadic kernel is the implementation and is to be mounted at or below #400,
so the architecture side of each conflict resolves to ours:

- src/_lcm/egm/upper_envelope/query.py   (--ours)
- src/_lcm/egm/nbegm_step.py             (--ours)
- tests/solution/test_envelope_query.py  (--ours)

The remaining two conflicted files are NOT architecture. `git log
4a53546..114b69b -- <file>` shows both carry only 9fae051, "Read the NB-EGM
continuation on each period's own grid" -- the cascade fix itself. Its change
there is additive fixture plumbing: optional `illiquid_grid` and `liquid_grid`
kwargs that the fix's own regression test passes. That test,
tests/solution/test_age_specialized_solver_composition.py, is not among the
conflicted files, so it lands here and calls both kwargs; under `--ours` it
would raise TypeError and the fix would arrive with its regression dead.

Both parents' edits to those two files are disjoint -- #406 adds the grid
kwargs, #407 adds outer_search/branch_aggregator/dead_functions -- so they
resolve as a union that drops nothing from either side:

- tests/test_models/n_nbegm_toy.py       (union)
- tests/test_models/nbegm_common.py      (union)

Verification on the merge result. Hooks do not re-run on a merge commit, so
`prek run --all-files` was run by hand: it caught 2x ruff FURB164 in
test_envelope_query.py -- #407 code newly linted under #406's ruff config, red
in neither parent -- fixed here. `ty` clean. `pixi install --locked` clean for
default and tests-cpu; `pixi lock` a no-op. Every line added by either parent is
present in the merge across all 14 both-touched files; the 111 theirs-only and
57 ours-only changed files are byte-identical to their source parent.

test_age_specialized_solver_composition.py 9 passed. The seven #400
certified-sign/double-double module test files 76 passed, 2 skipped;
certified_sign.py, double_double.py and cell_hull.py land intact and do not
depend on query.py, so both architectures coexist without a dangling import.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Brings across #400's `06bf038` (pin the envelope owner across eager, JIT and
vmap) and `a60a092` (bound a certified margin by its own residual), #390's
`af4b05b`/`ec96c12`/`ffb25e2`/`cdd4f65`, and the merges that carried them.

One conflict, and it is NOT architecture: `_build_Q_and_F_per_period` in
`src/_lcm/regime_building/processing.py` gained a keyword-only parameter on both
sides -- `reachable_targets` from #390's declared-reachability work, and
`period_to_regime_v_interp` from this branch. Both are needed and both are
keyword-only, so they union; nothing is dropped. Per-file derivation with
`git log 114b69b..18f5491 -- <file>` confirms no other conflicted file, and in
particular `query.py` did not re-conflict: the ownership resolution from the
previous merge holds.

`tests/solution/test_envelope_margin_bound_is_outward.py` is REMOVED here, and
this is the one judgement call in the merge. It arrives with `a60a092` and
imports `_value_quotient` from #400's `query.py` -- an internal of the envelope
architecture that the recorded three-branch resolution supersedes on this
branch, so the symbol does not exist and `ty` fails. The property it asserts,
that a certified margin's bound is outward and never understates its own error,
is NOT superseded and applies to this kernel too; `certified_quotient_margin`
is present here. Porting it against this branch's `_dd_div` is owed work, filed
in the fleet coordination directory rather than dropped.

Verification: no line added by either parent is absent from the result (the one
both-touched file checked line-by-line; 11 theirs-only and 66 ours-only files
byte-identical to their source parent); `prek run --all-files` clean, run by
hand because hooks do not re-run on a merge commit; `pixi install --locked`
clean for default and tests-cpu; `pixi lock` a no-op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Brings the 390->400->405->406->407 cascade to #409: ffb25e2, ec96c12, af4b05b,
06bf038, a60a092 and cdd4f65 are now all ancestors.

Both conflicted files carried real 407-side fixes, so neither side could simply
win and `-X ours` would have discarded them silently.

Stateless targets (ec96c12: a target carrying no state is worth its utility, not
zero) now seed `mixture_terms` instead of being accumulated as a raw `E += p*V`.
The incoming raw product would have reintroduced `0 * -inf = NaN` here, where
dissolution makes `-inf` continuations ordinary; routing them through
`_sum_regime_mixture` keeps the zero-mass masking and the value-ordered,
alpha-renaming-invariant reduction.

`_sum_regime_mixture` accordingly lifts rank-deficient values to the common value
shape before stacking, since a stateless target's V is rank-zero while a carry
target's contribution carries the cell and stakeholder axes. The `len < len` guard
covers the all-stateless case, where the values would otherwise right-align
against the probabilities and weight a cell axis by the target axis.

`reachable_targets` unions 407's declared-target keys with the existing per-target
and coarse-candidate sets rather than being replaced by them. Replacing would have
bypassed `_fail_if_coarse_candidate_folds_ambiguously` and the fold-round4 F2
guarantee. The union is safe in the direction that matters: a spurious candidate
carries zero transition probability and contributes nothing, whereas omitting a
routed one silently drops a continuation.

Restores the `QNAME_DELIMITER` import that 407 dropped along with its only use.

Verified: ty clean; tests/regime_building 368 passed; tests/simulation +
stateless-continuation 142 passed, 2 skipped, 6 xfailed. The one failure,
test_simulate_constraint_reads_next_state.py::test_the_constraint_is_available_as_an_additional_target,
is pre-existing and stack-wide -- it fails identically on 18f5491 (#406), 0e3a574
(#407 pre-cascade), f7aa590 (#407) and 28d05bb (#409 pre-merge), each checked in an
isolated worktree with _lcm.__file__ asserted inside the tree.
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Ten commits in one hop, held deliberately until #400's seven had reached #406
so this would not be two sequenced merges: `bf61b145` and its chain, `8f31991`
(publish a realized target reading a phase-invariant next state), `5b4339e`,
`d2f7b56`, `0b6be03`, `d9aec7d`, `3934701`, `f7ae72a`.

ZERO conflicts, which is where silent drops hide, so the per-file rule was
applied to the result rather than skipped: 16 theirs-only and 66 ours-only
changed files are byte-identical to their source parent, and all three
both-touched files (`regime_building/processing.py`, `solution/nbegm.py`,
`tests/simulation/test_simulate.py`) were checked line-by-line -- every line
added by either parent is present in the merge.

`model_processing.py` keeps the `from _lcm.processes.base import
_ContinuousStochasticProcess` form that #400->#405 settled on as strictly
narrower and cycle-safe; it did not resurface as a conflict here, and the module
was confirmed by importing it, not by reading it.

Verification: `prek run --all-files` clean, by hand, because hooks do not re-run
on a merge commit; `pixi install --locked` clean for default and tests-cpu;
`pixi lock` a no-op; 137 passed across the certificate guard, the cascade
regressions and `tests/regime_building`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Brings the combined cascade hop to #409: bf61b14 (#400), 8f31991 (#406
transition-dependence fix), ddbec5b and 5b4339e are now all ancestors.

One conflict, in `Q_and_F.py`, and both sides were right about different things.

#407 `d9aec7d` transforms a stateless target's value BEFORE the expectation:
`QuasiArithmeticMean` is `CE = g^-1(sum_r p_r * E_w[g(V'_r)])`, so `transform`
precedes every expectation including the regime-transition one, and `inverse`
applies once to the finished sum. Previously a stateless target was
probability-weighted raw and only then inverted, leaving the transformed domain.

#409 routes stateless targets through `mixture_terms` rather than a separate
`E += p*V`, so they get `zero_safe_weighted_term` (a zero-probability +-inf is
masked to exactly 0 instead of yielding NaN) and the value-ordered,
alpha-renaming-invariant reduction. On this branch dissolution makes `-inf`
continuations ordinary, so a raw product is a live NaN source, not a hypothetical.

Resolved by doing both: CE-transform the stateless value, then seed
`mixture_terms` with the transformed value. `_sum_regime_mixture` forms
`sum_r p_r * g(V_r)` zero-safely and `ce.inverse` is applied once to the result,
which is exactly `g^-1(sum_r p_r * g(V'_r))`.

Also drops `test_collective_solve_default_call_still_returns_bare_mapping`
(external test-suite dedup audit, round 1): it is a local mirror of
`test_collective_regime_simulate.py::test_public_model_solve_default_return_shape_is_byte_identical`,
which is the module it already imports its model helper from, and whose
assertions are a strict superset. The file's own singleton guard,
`test_singleton_solve_bare_call_returns_legacy_mapping_not_a_tuple`, is retained
-- that is the invariant this module exists for. The other five deletions from
that audit are shared with `main` and went to #390 as a patch.

Verified: ty clean; tests/regime_building + the two touched test files 379 passed;
the certainty-equivalent bearing files (test_temporal_aggregation,
test_dcegm_validation, test_nbegm_epstein_zin_composite_flow,
test_simulate_terminal_rows) 43 passed.

KNOWN COVERAGE GAP, not closed here: no test appears to construct a model with
BOTH a certainty equivalent AND a stateless target, so the intersection this
resolution creates is exercised by neither side's tests. `d9aec7d` itself shipped
without tests. The two halves are covered separately; their combination is not.
hmgaudecker added a commit that referenced this pull request Aug 4, 2026
Carries nb-egm (#400), state-conditioned shocks (#405) and perceived stochastic
transitions (#406) through #407. Two conflicted files, 15 hunks, and the
resolution is a PORT rather than a side-pick: the incoming branches refactored
exactly the machinery this branch had extended.

KNOWN RED: 13 tests still fail, all one class, see the last section.

## Q_and_F.py (5 hunks)

The incoming extraction `_get_compute_E_next_V` is adopted at BOTH call sites
and this branch's two inline aggregation blocks are gone -- but their numerics
were ported INTO the helper, not dropped:

- Linear path: `zero_safe_average` within a target, and the cross-target mixture
  reduced ONCE by `_sum_regime_mixture` -- stack the operands, one zero-safe
  contraction, value-ordered sum. That preserves round-8 (accuracy: stacking the
  operands rather than the formed products lands on the exact-policy side of the
  pinned 5-target fixture) and round-10 F1 (the reduction order is a function of
  the contribution multiset, never of regime LABELS, so an alpha-renaming cannot
  move the result). Upstream accumulated `E + p*V` in label order, which would
  have reintroduced both defects -- and `0 * -inf = nan` besides, which is
  ordinary here because dissolution makes `-inf` continuations routine.
- `_scalar_target_contribution` now returns UNMULTIPLIED terms, so a stateless
  target joins that same single contraction instead of being summed separately.
- CE path keeps UPSTREAM's single joint `ce.aggregate`. That deliberately
  supersedes this branch's per-target `transform -> reduce -> inverse`, which
  cannot anchor `PowerMean`; flattening before aggregating is the whole point of
  the incoming fix.

`get_period_targets` and `get_period_scalar_targets` are DELETED. They had no
live callers left after the merge, and #400 ships guards that forbid them by
name (`test_engine_has_no_period_target_inference_helper`) and ban the literal
`carry_targets` (`test_continuation_targets_are_not_derived_from_law_bundle_keys`).
Both guards pass. Comments that named them were reworded rather than left to
resurrect the banned vocabulary.

Git had also FUSED two valid states into a broken one here: it aligned
`def _get_compute_E_next_V(` with this branch's `get_period_targets` tail on the
two lines they share, splicing one function's body inside the other's signature.
Reconstructed deliberately rather than patched in place.

## processing.py (10 hunks)

Mostly unions -- the incoming `phase_reachability`/`source_regime_name` and this
branch's `coarse_state_law_names`/`fold_only_regimes` are different parameters
serving different concerns, and both are read in the merged body. Substantive
ones:

- `reachable_targets` (this branch, still read by the fold-only empty-bundle
  rule) and `continuation_targets` (incoming, the graph-retained process
  scoping) are BOTH kept: different sets, both live. The coarse-candidate
  `_fail_if_coarse_candidate_folds_ambiguously` guard is preserved.
- The collective builder takes `period_targets=stateful_targets` from
  `partition_continuation_targets`, which is exactly what this branch's
  `period_targets` has always meant.
- Removed a duplicate `source_regime_name` the union introduced at two call
  sites and in `_process_regime_core`'s signature (caught by `ty` as a genuine
  `invalid-syntax` duplicate keyword).

## Tests

38 `process_regimes` call sites migrated to #400's now-required
`prepared_structure`, via `tests.conftest.build_prepared_structure`, including
the `_solve_kwargs` helpers in `test_fold_gate_guard.py` and
`test_fold_guard_complete.py`.

## KNOWN RED: 13 failures, one class, remedy already decided

#400's new build-time handoff check (`2bfc333`, adopted by `94e868e`) rejects a
target that declares a stochastic process the source neither carries nor
supplies an entry law for. This branch's gated-edge / fold models exercise
exactly that shape -- the deleted `get_period_targets` docstring named it a
"known remaining hole (fold-review F2, deliberately NOT closed in this slice)",
and #400 closed it by rejecting such models.

Affected: `test_fold_iid_shocks.py`, `test_gated_edge_arg_provenance.py`,
`test_fold_guard_complete.py`, `test_fold_gate_guard.py` (31 instances of the
same rejection).

Decided remedy (maintainer): fix the MODELS, not the check -- give each source a
target-specific ENTRY LAW, `state_transitions={"wage_shock": {"target": <law>}}`
(shape already used at `test_fold_iid_shocks.py:471`). An entry law leaves the
source's state space, and so the fold/gate topology under test, unchanged.
Applied in the follow-up rather than here, because `feat/nb-egm` has moved
(`4cb10b6` -> `fdd5d6d`) and this merge is redone against the new tip anyway.

Verified: `prek run --all-files` green (ruff, ruff-format, ty); 597 of 612 pass
across tests/regime_building/, test_Q_and_F, test_collective_regimes_draft and
test_certainty_equivalent. Not the full suite. NOT pushed: the branch is
knowingly red on the class above.
hmgaudecker added a commit that referenced this pull request Aug 4, 2026
Clean merge, no conflicts: nb-egm's four new commits reach this branch through
#405/#406/#407, touching five test files only. None of the step-4 port recorded
in 8ab639e is disturbed.

prek --all-files green (ruff, ruff-format, ty). The branch is still red on the
target-only-handoff class documented in 8ab639e; the fix follows.
hmgaudecker added a commit that referenced this pull request Aug 5, 2026
Completes the cascade #405 -> #406 -> #407 -> #409, carrying `feat/nb-egm`
@ `48d3795c` and with it the aggregator specification objects, the EGM
preference-map refactor, and the runtime inactive-target check.

Fourteen conflicts, and unlike the branches above this one they are not all
import lists.

`W_linear` is gone, replaced by the `LinearAggregator` dataclass. This
branch had 88 references across 14 test modules, all of the form
`koopmans_aggregator=W_linear`, plus one docstring. Migrated to
`LinearAggregator()`. These would have failed loudly at import, so they were
never the dangerous half; the dangerous half is the identity check that
silently stops gating, and there is none left -- no `W_linear` or
`W_epstein_zin` survives in `src/` or `tests/`.

Certainty-equivalent reduction (`Q_and_F.py`), the substantive one. This
branch collects UNMULTIPLIED `(regime, p, E[V])` terms and contracts them
ONCE through `_sum_regime_mixture` -- value-ordered, and deliberately not
keyed on regime labels (round-8 accuracy, round-10 F1). Upstream folds
`CE = CE + p*E[V]` per target and divides by the represented mass. Kept both:
the single contraction produces `CE`, then upstream's mass gating applies to
it unchanged (`_unit_regime_mass_or_nan` on the per-target route,
`_regime_mass_is_unit` selecting against NaN on the lottery route). These
compose exactly because `_sum_regime_mixture([])` returns `zeros_like(like)`,
which is the zero upstream's accumulator started from.
`_scalar_target_contribution` now returns the mixture terms AND the
probability mass, rather than one or the other.

The reducer stays `zero_safe_average`, NOT upstream's new
`_expectation_over_stochastic_nodes`. The latter guards the weight SUM only,
so `0 * -inf` still poisons its numerator -- and `-inf` continuations are
routine here, where dissolution makes them ordinary rather than exotic. The
PREDICATE does move to upstream's `continuation.has_lottery_axes`, since the
`next_V_has_stochastic_states` mapping it replaced no longer exists and would
have been a `NameError`.

`_get_joint_weights_function` changed contract upstream: it takes the ordered
tuple of lottery variables instead of re-deriving them, so that the weight
axes and the productmapped value-surface axes are fixed by one ordering.
Upstream updated the singleton path only; the collective builder is this
branch's and still called the old signature. `ty` caught it. Updated to
derive `lottery_variables` exactly as `_build_target_continuation` does.

Duplicate-parameter trap, three times: `source_regime_name` is passed at the
top of both `_process_regime_core` call sites and appears again in upstream's
added block, and is declared twice in the signature -- which is a
`SyntaxError`, not a type error. Kept one of each and took only `phase_name`.

`ISC004` fires on six of this branch's error strings under the ruff config
that arrived with the merge; wrapped each in parentheses rather than
silencing it, since every one sits in a returned collection where a dropped
comma would fuse two distinct messages. `PLR0915` then fired on
`_process_regime_core` (55 > 50) because the merge added statements;
extracted `_partition_phase_functions`, which is a pure move -- `all_functions`
and the stochastic-name set were already local to that block.

One real defect surfaced, the known coarse-law witness from
`BUG-gridsearch-drops-inactive-target-mass.md`:
`test_coarse_self_transition_retains_the_self_continuation` routed its entire
mass to `stay` at `stay`'s last active age, where `stay` is inactive next
period. It solved to an all-NaN value function and said nothing, because the
test ran at `log_level="off"` where the validation is skipped. The coarse law
is now age-dependent and routes to `done` at that age; period 0 -> 1 still
exercises the F1 regression this test exists for, namely that a coarse law
returning its OWN regime keeps the self-continuation. Converted to
`log_level="debug"` per the fleet NOTICE, so the check actually runs.

748 passed across `tests/regime_building/` plus the aggregator tests, none on
the AVOID list, `_lcm` asserted to this worktree on the controller and both
xdist workers. `prek run --all-files` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 17, 2026
Leg 4 (final) of the dcegm review-repair cascade. Seven conflicted files.
Four were unions of two independent additions; three needed a decision.

Mechanical unions (both sides kept):

- `.pre-commit-config.yaml` — `name-tests-test` exclude is now the three-way
  union: `tests/mock_regime.py|tests/collective_fixtures.py|
  tests/envelope_configs.py|^tests/test_models/|^tests/data/|^tests/solution/_`.
  Both fixture modules exist in the merged tree.
- `src/_lcm/model_processing.py`, `src/lcm/__init__.py` — import and `__all__`
  collisions from alphabetical interleaving. `ExactEnvelope` was already
  imported at line 165, so the `__all__` entry it gained is backed.
- `src/lcm/model.py` — `_solve_compiled` now takes BOTH
  `collect_simulation_policies` (from #407) and `retain_dissolution_flags`
  (from #409); its body already called `solve()` with both. The auto-solve
  comment block keeps #409's dissolution-flag note and #407's two-signal note,
  since they document different things.

Decisions:

1. `src/_lcm/simulation/simulate.py` — the two branches widened the SAME
   function differently. #409 returned `(actions, nested_fallback)`; #407
   returned `(actions, reported_value, nested_fallback)`. The auto-merged
   return paths had already taken #407's 3-tuple, so the conflicted paths were
   resolved to match it, and the docstring to #407's three-item form (#409's
   fallback description is a subset of it). Verified afterwards: the signature
   and all six return paths are 3-tuples.

2. `src/_lcm/solution/backward_induction.py` — #407 recomputed
   `base_state_action_spaces` inside `solve()` at the old position while #409
   had MOVED that computation earlier in the same function, so that a
   `_reject_edge_fold_state_param_collisions` fence could run before any kernel
   is compiled. Textually these are additions at different offsets, so the
   merge offered both and neither side conflicts — the relocated-refactor
   duplicate. #407's copy is dropped; #409's earlier one is kept, which
   preserves the fence's position. Kept #407's `host_device` gating
   (`None` unless policies are collected) and merged the two comments.

3. `src/_lcm/regime_building/processing.py` (7 hunks) — #407 renamed
   `is_egm_carry_target` → `continuation_demanded` and replaced the inline
   `egm_child_regimes` comprehension with the `_continuation_targets` helper.
   Took the rename at all six sites and the helper call, while keeping #409's
   four blocks in the same region (`fold_only_regimes`, the same-period-ref
   validation pair, the gated-edge validation, the folded-endpoint check),
   which #407 has no counterpart for. `is_egm_carry_target` and
   `egm_child_regimes` now appear nowhere in `src/` or `tests/`.
   `solver_kernels.continuation_template` → `.continuation_spec` came with the
   same wave; `continuation_template` survives as the delegating property on
   `engine.py:377`, so the two forms are equivalent.

Verification: `ty` green; all six resolved modules parse; no conflict markers
remain anywhere in the tree. Behavioural smoke, all on CPU:
- `tests/simulation/test_egm_sim_policy_value_read.py` +
  `test_egm_associated_policy_read.py` — 5 passed in 1.20s
- `tests/simulation/test_simulate_n_nbegm_continuous.py` — 5 passed in 146.02s
  (the file whose four tests the #407 leg broke and repaired)
- `tests/solution/test_collective_gated_edge_age_specialized_axes.py` +
  `tests/simulation/test_aot_collective_and_gated.py` — 7 passed in 15.02s
  (this branch's own feature, the part decisions 2 and 3 could have broken)

Known gap, cascaded in from upstream and NOT introduced here: this wave
predates nb-egm `4db23f8b`, so all four legs carry the Windows-only
exact-kernel break — `ModelInitializationError` raised where the conftest skip
gate keys on `ExactAffineKernelUnavailableError`, 32 tests red on Windows only.
The fix cascades as the next wave from nb-egm `dfa68097`.
@codecov

codecov Bot commented Aug 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.51633% with 174 lines in your changes missing coverage. Please review.
✅ Project coverage is 92.67%. Comparing base (9987966) to head (1c9f1d2).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/lcm_examples/mahler_yum_2024/implicit_pilot.py 41.00% 82 Missing ⚠️
src/_lcm/simulation/simulate.py 91.76% 21 Missing ⚠️
src/_lcm/solution/nnbegm.py 93.20% 18 Missing ⚠️
src/_lcm/egm/outer_affine_structure.py 90.35% 11 Missing ⚠️
src/_lcm/regime_building/processing.py 76.92% 9 Missing ⚠️
src/_lcm/egm/outer_refinement.py 96.84% 6 Missing ⚠️
src/lcm_examples/mahler_yum_2024/paper.py 89.83% 6 Missing ⚠️
src/_lcm/egm/outer_candidates.py 86.84% 5 Missing ⚠️
src/_lcm/egm/outer_carry.py 94.02% 4 Missing ⚠️
src/_lcm/egm/nested_published_policy.py 95.38% 3 Missing ⚠️
... and 4 more
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #407      +/-   ##
==========================================
- Coverage   92.86%   92.67%   -0.19%     
==========================================
  Files         285      310      +25     
  Lines       29951    32809    +2858     
==========================================
+ Hits        27814    30407    +2593     
- Misses       2137     2402     +265     
Flag Coverage Δ
cpu-python 92.67% <91.51%> (-0.19%) ⬇️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

hmgaudecker and others added 8 commits August 19, 2026 18:50
A regime's slot type decides the base: a plain Regime takes any Solver, a
ConsumptionSavingsRegime takes a OneMarginSolver, and a nested one takes a
TwoMarginSolver. Each specialised regime checks this at construction, so
subclassing Solver directly cannot produce a solver they will accept.

Adds the margin families, why the regime owns the DAG role names while the
solver carries numerical configuration only, how to implement the binding
method through bind_roles so a subclass reaches the engine as itself, and
the worked example in the tree for each of the three bases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yHhuFzqsDw1MhdB1i2Ljm
…tions' into feat/continuous-outer

# Conflicts:
#	tests/test_consumption_savings_regime.py
The base-class table read as though a plain `Regime` takes any `Solver`. It
does not: `_fail_if_egm_solver_has_no_margin_declaration` rejects a
`OneMarginSolver` or `TwoMarginSolver` handed to a plain `Regime`, because the
four role names such a solver needs have nowhere to be declared. That is the
one row a reader consults to pick a base class, so it has to name the exclusion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
Both nested toys reached the next durable stock through an investment
action, so the oracle scored a candidate set the nested solvers never
sweep. With `Z = 20k/9` on the `illiquid` grid, `s' = 20j/14` on
`OUTER_GRID` and a unit investment step, the two sets coincide only where
63 divides `9j - 14k`: 3 of the 15 outer nodes, and those reachable from
only 2 of the 10 `illiquid` states. The other 8 states hit none at all.
A comment above `ILLIQUID_INVESTMENT_GRID` asserted the opposite, which is
why the gap read as tolerance rather than as a defect.

Both toys now let the brute variant choose `new_illiquid` directly on
`OUTER_GRID`. The `illiquid` grid is untouched, so the state space is the
same before and after and the two runs are comparable.

Checked against an enumeration written outside pylcm, continuing through
the same interpolated terminal carry the solvers use: the repaired brute
arm reproduces it to 0.0000 at every probed cell, so it is now an oracle
rather than a second approximate solver.

That makes the comparison directional, and the test says so: on the
interior the nested solve dominates in all 72 cells, max relative gap
0.0176, identical at both precisions, where the old unsigned band was
0.25 and measured nothing.

The poorest wealth row is excluded rather than loosened. There the exact
enumeration puts the attainable optimum at (wealth 0, illiquid 0) 6.7
utils below what NNBEGM reports and 6.8 below NEGM, so it is the inner
families' constrained-region representation, shared by both and outside
what this file pins. It is documented in the docstring and asserted
nowhere.

(cherry picked from commit 37d6f19)
…tions' into feat/continuous-outer

# Conflicts:
#	tests/solution/test_n_nbegm_matched_oracle.py
#	tests/test_models/n_nbegm_toy.py
A regime accepts any `OneMarginSolver` or `TwoMarginSolver` instance, but three
engine sites dispatch on the concrete shipped classes: the simulate-phase
budget-constraint synthesis, the terminal branch of the budget constraint's
inner lookup, and the carry-target read. A solver deriving straight from a
marker was therefore accepted, bound, and then solved with the budget
constraint and the carry read skipped — a wrong published policy, with nothing
raised.

`fail_if_solver_is_not_shipped` refuses such a solver at model build, from both
`validate_model` loops, so no custom solver reaches any of the three sites. A
subclass of a shipped solver stays supported: it inherits the concrete type
every site tests for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
hmgaudecker and others added 16 commits August 27, 2026 14:21
Recovering the outer map's slope by differencing two rounded evaluations
is not accurate enough to invert with: for `target = state + action` the
true slope is exactly one, but the secant returns ten distinct values over
a grid sweep, and dividing by one of them carries the reassembled stock
off the outer state's declared domain.

Read the coefficient off the traced structure instead, in exact rational
arithmetic. A forward walk over the jaxpr seeds the outer action's
variable with coefficient one and propagates it through the affine
primitives; any primitive that consumes a tainted variable without an
exact affine rule refuses the map and names itself.

Automatic differentiation cannot do this job, which is why the walk reads
arithmetic rather than derivatives: `stop_gradient` is invisible to AD, so
`state + action + stop_gradient(2 * action)` linearizes as unit slope
while the map moves three units of stock per unit of action. The walk
returns three and refuses to call it unit.

The module is standalone and not yet wired into either inversion path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Both phases recover the outer action from a retained target, and they must
recover it the same way: a solve that certifies a candidate and a
simulation that replays a different action for it disagree about which
stock was chosen. Give them one inverse and one admissibility predicate.

The inverse takes the coefficient from the declared structure rather than
from a two-point probe, and elides the division entirely at unit slope, so
the rounding the estimated-slope route introduced is gone rather than
bounded. A non-dyadic coefficient rounds its quotient however exactly the
factor is known, so it is refused where it is declared; a dyadic one is
necessary structural support and never sufficient pointwise certification,
so the runtime predicate stays active for it too.

Admission is exact and two-tier. Every image must lie inside the outer
state's declared domain, because a stock outside it has no value function.
An image whose target is a declared endpoint must reproduce that endpoint
bit for bit, because that is where an ULP of error leaves the domain. No
scale constant appears anywhere on the acceptance path.

This fixes what a represented interior candidate means, and states the
weaker meaning plainly: an inverse-produced `(action, image)` pair, not a
claim of bit-exact target identity. An interior round-trip difference is
an approximation diagnostic.

One limit is pinned rather than assumed: XLA flushes subnormals, so the
exact comparison resolves a zero endpoint to the smallest normal
magnitude, not to the last subnormal.

The module is not yet wired into either inversion path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
N-NB-EGM searches the outer margin over post-decision targets and stores
those, so both phases must recover the action that reaches a stored target.
Both estimated the map's slope by differencing two rounded evaluations and
divided by the estimate. For the supported law `new = old + action` the
secant is one ULP from one, and dividing by it carries the reassembled stock
off the outer state's declared domain -- where there is no value function --
while a scale-relative round-trip tolerance accepted the result.

Both phases now read the coefficient off the traced structure of the declared
map in exact rational arithmetic and share one inverse. Only a constant
power-of-two coefficient divides exactly, so anything else is refused where it
is declared rather than left to miss a node per candidate. The adaptive
path's numeric unit-slope check is the same certificate, so its tolerance
constants are gone too.

Admission is tolerance-free and two-tier: every candidate's image must lie
inside the outer state's declared domain, and a candidate whose target is a
declared endpoint must reproduce it bit for bit. Interior candidates are
admitted on containment, which fixes what a represented candidate means -- an
inverse-produced (action, image) pair, not a claim of exact target identity.

A candidate that fails is dropped, never projected inward: a projected pair is
not the pair the solve ranked, and rescoring cannot undo pairing a moved outer
action with an inner action obtained for the nominal candidate. Simulation
drops pointwise and announces the aggregate per regime-period; the solve, whose
candidates all sit at declared grid nodes, stops instead -- a drop there means
the law and the grids disagree. Reading either count back is a host transfer,
so the solve pays it only in raise mode, matching the backward-induction
NaN fail-fast.

Replay also takes an adjuster candidate's target from the declared search
nodes rather than the interpolated bank. That surface is constant along the
state axes, yet linear interpolation forms `(1 - w) * c + w * c`: measured over
100001 coordinates it misses node 20.0 for 754 of them, by up to 16 float32
ULP, and a target a rounding away from an endpoint is not that endpoint.

A state-dependent slope is no longer supported and is refused at build. The
test that asserted such a law inverts correctly encoded the divide-by-the-
estimated-slope derivation being removed, and now asserts the refusal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
A candidate whose target is a declared node can indict the declaration: the
outer search grid and the keeper's own grid value are fixed by the model, so
an inverse leaving the outer state's domain there means the declared law and
the declared grids disagree. That is worth stopping a debug-level solve on.

A regime that declares its own no-adjustment target does not produce such a
candidate. It computes the keeper target from the DAG at each solve cell, so
the value can sit off-node for reasons that are economically ordinary -- a
depreciating durable reaches a stock no grid node names. That is the same
category of thing simulation meets at a realized state, and it is dropped the
same way rather than treated as a contradiction.

The refusals now name the solver that solves what N-NB-EGM cannot. Telling a
modeller to "choose a solver that does not invert the outer margin" states a
property and leaves them to find which solver has it; `GridSearch` searches
the outer action directly, so it never inverts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
A domain endpoint is a node value, so it belongs to the period being
solved. Both phases captured it once from the representative age's grid
instead, which is the wrong domain everywhere else under an
age-specialized outer grid: a stock only the later ages hold was judged
out of domain and dropped, and a stock past a narrower age's edge was
admitted with no value function to read it on.

Solve resolves the endpoints per period from `period_to_state_nodes`;
simulation reads the same per-period axes off `period_state_axes` rather
than forming its own from the published grid, so the two phases cannot
disagree about which stocks are representable -- the divergence the
outer-target inversion exists to prevent. Both fall back to the
published grid for an age-invariant regime, where it is the whole story.

Measured on the toy: 36 of 1829 candidates were dropped before, 0 after.

Also retires comments describing a two-point slope probe that the
structural coefficient certificate replaced, and narrows the
age-specialized composition test's outer search to the stocks every age
holds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Solve reads a period's age-specialized node arrays from
`SolverBuildContext.period_to_state_nodes` and simulation from
`SolutionPhase.period_state_axes`. The two are built by different
functions from different inputs, with different key-set expressions,
and both fall back to the published grid for a period they do not
carry -- so a disagreement between them would not surface as an error.
It would surface as the solve certifying candidates against one set of
domain endpoints while simulation replays against another, which is the
divergence the outer-target inversion exists to prevent.

Compare the two tables where they are built and refuse a mismatch,
naming the period, the state, and the first index at which the node
arrays part. The comparison is exact: these are node values that must
be identical, and a tolerance wide enough to absorb a regridding is
wide enough to hide one. Key sets must match exactly rather than by
containment, so the guard does not depend on the current derivation of
one from the other continuing to hold.

An age-invariant regime resolves `None` in both phases and is
unaffected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Two class-soundness defects in the round-3 outer-target inversion, both
found by the closure audit's reviewer-executed oracles.

`convert_element_type` was a blanket pass-through in the structural
coefficient walk, so a map that round-trips the action through a
narrower dtype certified slope one. float32->float16->float32,
->bfloat16-> and ->int32-> all certified, and containment-only interior
admission then accepted their wrong images: a target near ten comes back
3e-3 low through float16 and 3e-1 low through int32, which is not a
rounding and is still inside the declared domain, so containment cannot
catch it. The round-2 residual gate rejected these; the exact route
accepted them, because the structure is affine while the arithmetic
quantizes. The primitive is refused outright rather than carried with a
per-dtype rule: no supported map converts the action, and a widening
conversion, though value-preserving, is not worth a branch. Traced
jaxprs confirm ordinary maps emit no convert_element_type at all, so
this refuses exactly the lossy form.

A declared solve-node that cannot be reconstructed also raised only at
log_level="debug" and otherwise published a NaN-reduced candidate bank.
The failure is known before publication, so gating it on the log level
made the published policy depend on the diagnostic setting -- the same
model publishing one policy at "off" and another at "debug". It now
refuses at every level, and logs before raising so a caught exception
still leaves a record. A failure first met at a realized off-grid
subject keeps drop-and-announce: there the state landed awkwardly, which
the model author cannot be expected to have precluded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The certificate refuses `convert_element_type` because the primitive
preserves an array's shape and not the action's value. Three cases that
rule has to survive, none of which the lossy round trips covered:

- a widening conversion, which is arithmetically safe and refused
  anyway, so the rule stays a property of the primitive rather than a
  per-dtype exactness table nobody can keep right;
- a conversion hidden inside a `jit` wrapper, which is refused only
  because the walk descends into sub-jaxprs -- had descent stopped at
  the call primitive the wrapper would certify slope one, the same false
  certificate with a layer around it;
- the pass-through set itself, asserted directly so restoring the
  primitive fails on the rule and not only through behaviour.

Restoring `convert_element_type` to the set kills seven of these tests,
measured, so the guard is known to fire rather than assumed to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
A dropped candidate carries `nan` in the published value surface, so a
bank with every candidate dropped masks every objective to `-inf` and
`argmax` over them returns index 0 by convention. A reduction that
trusted that index would emit candidate zero as the winner -- an action
the solve never represented.

The fail-closed sentinels already survive; nothing pinned that they do.
The existing empty-valid-set test exercises a different input path:
represented candidates that score invalidly, not candidates that were
never represented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
A solve that cannot reconstruct a declared node refuses at every log
level, so both workflows must surface that refusal identically: the
split one calls solve directly, the automatic one calls it from
simulate when no value arrays are supplied. Without a test, a refusal
reaching one route could be bypassed in the other.

The witness is a signed outer domain. Recovering the action as
`target - at_zero` and adding it back crosses zero there, so some
declared nodes are reached only to a stock outside the domain. Matched
positive-domain grids solve, which is what attributes the refusal to
the domain's sign rather than to grid size or the outer search.

Also state the bound rather than one sample point in the
`convert_element_type` comment: an int32 hop discards the whole
fraction, so it moves a target by up to a unit, not by the 3e-1 that
happens to hold at 10.3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The adaptive mesh published a replay policy for a map nothing could
invert. Certification lived inside `_candidate_outer_targets`, which
only the finite search calls, so a declared outer map routing the outer
action through a lossy conversion was refused on the finite grid and
accepted on the mesh. Simulation then replayed from a policy whose
outer action cannot be recovered, selecting a different decision
without saying so.

The certificate belongs to the declared map, which both searches read
and neither owns, so it is resolved once per period in the dispatcher
before either can run. It reads structure -- argument names and shapes,
never values -- so scalar stand-ins certify the same map the search
later binds to full candidate arrays. The finite route consumes that
one object rather than resolving a second; the adaptive route is gated
by the resolution but is not handed it, because it never inverts.

The toy resolves its outer post-decision override at call time rather
than capturing it as a default argument, so `new_illiquid` stays
patchable: a default binds at definition, which silently stopped three
refusal tests from refusing and left a fourth passing while exercising
nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The outer certificate admits any power-of-two coefficient, but the
continuous-outer reader recovers the action by subtracting the declared
map's offset, which inverts only at one unit of stock per unit of action.
A doubled map therefore certified, published a nested policy, and failed
later in simulation as an infeasibility -- naming neither the inverse nor
the coefficient.

Thread the per-period certificate into the adaptive search and gate
publication on a unit coefficient, so the refusal lands before any policy
object exists and names the map. A regime that publishes no nested policy
is unaffected.
The capability question -- can the declared outer map be inverted, and can
its post-decision and keeper candidates be resolved against the simulate
pool -- was answered in three places: at finite publication, again by the
nested reader re-tracing the user transition, and again at the finite
replay's own re-certification. Simulation could therefore reach a different
answer from the solve, and where it did the reader silently substituted the
generic grid-argmax pair for a whole regime.

Resolve it once per regime-period, before either search runs, and carry it
on both policy payloads as a required field. Every structural refusal now
raises before a policy object exists; simulation consumes the settled
verdict instead of rediscovering one, and keeps only the pointwise branch
-- residual, support, finiteness, budget -- that a realized state can
decide. Bindable names come from the params template, so one certificate is
authoritative for every call.

An adaptive regime with more than one passive axis is refused rather than
silently downgraded; FiniteOuterGrid replays it.
Brings in the taste-shock AOT compilation fix (#429). It touches
src/_lcm/simulation/compile.py and its own tests only, with no overlap
against the outer replay capability work on this branch.
Adaptive replay decided pointwise admission with a scale-relative residual
tolerance while finite replay used the tolerance-free predicate. The two
therefore disagreed about a declared endpoint: a recovered action landing one
ULP inside the domain reproduced the endpoint to within the tolerance but not
bit-for-bit, so adaptive published an action whose image sat beside the ranked
endpoint and finite refused the same candidate.

Both routes now call `outer_candidate_is_admissible`, so a declared endpoint is
represented only when the recovered action reproduces it exactly, and an
interior target only has to land inside the declared domain. The domain
endpoints are cast to the image dtype at the call site, matching what
`invert_declared_outer_target` already does: they are stored as Python floats
and the predicate takes scalar arrays.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The N-NB-EGM outer replay binds the user's outer target function by
merging the regime's flat params with the state, action, period and age
inputs. Params are not all float arrays -- the mapping holds
`FloatND | IntND | BoolND | MappingLeaf | SequenceLeaf` -- so declaring
the binding as `dict[str, FloatND]` claimed more than it delivers, and
the claim is enforced at runtime by the beartype claw. `EconFunctionArg`
is the alias for exactly this vocabulary: one argument of a user economic
function, as `EconFunction.__call__` accepts it.

Hook revisions move with it, and the ty bump is what surfaced the
mismatch:

- pyproject-fmt v2.28.0 -> v2.28.2
- ruff-pre-commit v0.16.3 -> v0.16.5
- ty-pre-commit v0.0.72 -> v0.0.75

`pixi.lock` is regenerated from scratch. Its platform table now names the
CUDA variants (`linux-64-cuda12`, `linux-64-cuda13`,
`linux-aarch64-cuda13`) instead of carrying them as `p1`/`p2`/`p3`; the
platform set itself is unchanged at seven.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
hmgaudecker and others added 8 commits August 29, 2026 10:11
The adaptive reader refuses a policy proposal whose recovered action lands
beside a declared endpoint rather than on it. Control then reaches the grid
baseline, which projects the raw pair onto each published mesh endpoint and
admitted the result on domain containment alone. An action one representable
step inside an endpoint satisfies containment, so the projection could publish
the very stock the proposal had just refused, and a signed outer domain reaches
that case for every subject whose stock and endpoint straddle zero.

Each projection now reports whether its action reproduces the endpoint it aimed
at, decided by the same tolerance-free predicate the proposal uses, and the
baseline may rank a projection only where it does. The predicate reads the
published mesh rather than the outer state's declared bounds, so a mesh that
stops short of those bounds still has endpoints the check can recognise.

The retreat to an endpoint's interior neighbour goes with it. It served
subjects whose exact projection left the mesh, but it did so by aiming at a
different stock from the endpoint being certified, which is the distinction the
verdict now draws. The two endpoints are mirror images, so a subject missing
the near one keeps the far one, which a sign-crossing action reaches exactly;
a subject missing both leaves the read with nothing to publish and raises.

The two-asset toy gains a terminal-utility override, since its default bequest
is singular one unit into durable debt and a signed outer domain needs one that
is finite there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
A rejected policy replay handed the fallback a choice between the two
declared endpoints and nothing else. The published policy carries far more
than that -- a keeper branch plus one conditional inner policy per outer mesh
node -- so a subject whose recovered action could not reproduce the endpoint
the solve ranked first was sent to the opposite endpoint even when a mesh node
between them was both reachable and better. On the shipped three-node mesh
`[-1, 4.5, 10]` that emits the far endpoint at canonical Q -9 while the middle
node, which the subject reaches exactly, is worth -3.5.

The fallback now ranks every branch the solve published. Each is admitted by
the same tolerance-free predicate the on-mesh proposal uses, against the same
declared inverse domain, so an unreachable branch removes that branch and
nothing else; each surviving branch keeps its own inner action rather than
another branch's; and the ranking is by the values the solve published, which
is the comparison the solve itself made across these branches. The keeper
occupies the leading row so it wins ties, matching the on-mesh convention that
an exact keeper beats an equal-valued adjustment. The caller scores only the
winner through canonical Q, which is what makes it comparable with the raw
simulation-grid pair.

Ranking a read whose leading axis is not the outer-node axis would pair every
node with another node's policy, so the length disagreement is refused rather
than ranked.

The endpoint-only selector is deleted; nothing projects onto an endpoint any
more. A subject reaches the refusal only when no branch at all is reachable,
which needs a mesh whose every node is a declared endpoint and a keeper target
outside the published domain.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The read hands its caller the branch a rejected replay falls back to, so the
return description names it and says what the caller may do with it. The
caller's own summary of the value it reports drops the word "projected": a
rejected replay is given a published branch, not a pair projected onto an
endpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Two changes that the pre-commit config forces into one commit: a hook rejects
any commit made while `.pre-commit-config.yaml` is modified but unstaged, so
the workflow bumps cannot land separately from the hook that gates them.

Hooks. Five the config was missing:

- `pixi-lock-check` at pre-push, so a stale lock fails here instead of burning
  a CI cycle. `--check` re-solves without writing, so it cannot dirty the tree.
- `pre-push-hooks-installed`, which turns a clone that never ran
  `prek install -t pre-push` from a silent no-op into a visible failure.
- `notebook-cell-source-format`, `no-hardcoded-user-paths` and `codespell`.

The first two currently find nothing, which is the point of installing them
before they would.

codespell reported 103 hits and every one is correct as written. `statics`
(comparative statics) accounts for 91, `arithmetics` and `disjointness` for
another six; those are jargon and go in the dictionary, which now matches the
shared one. Four are project-specific and take an inline ignore at their site
so the word stays caught everywhere else: the MSVC object-output flag, which has
to be spelled exactly; a deliberately misspelled key in the test that asserts
unknown keys are rejected; and a regex prefix that has to match both singular
and plural. Two prose uses of "re-uses" are simply reworded.

CI versions. `pixi-version` was pinned to v0.76.1 while the pixi that writes
`pixi.lock` locally is 0.78.0, so CI was validating a lock produced by a newer
binary with an older one. pixi is largely forward-tolerant, so that fails
quietly rather than loudly, and only stops working once the lock uses something
the pinned version cannot read.

`actions/upload-artifact` was three majors behind. v5 and v6 move the runtime to
Node 24 and require an Actions Runner of at least 2.327.1; the five ubuntu-latest
call sites are unaffected, and the two on `ubuntu-latest-gpu` will tell us on the
next run whether that runner is current. `prefix-dev/setup-pixi` goes to v0.10.2.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
`pre-push-hooks-installed` and `notebook-cell-source-format` run scripts from
the `.ai-instructions` submodule. pre-commit.ci clones without submodules, so
both fail there with "Executable /code/.ai-instructions/hooks/... not found"
rather than running. No workflow in `.github/workflows/` checks out submodules
either, so these two can only ever run in a developer clone.

Skipping them states that plainly instead of leaving two hooks permanently red.
The first asks whether this clone installed its pre-push stage, which a runner
has no stake in. The second genuinely loses CI coverage: notebook cell source
formatting is now checked where the submodule exists and nowhere else.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Neither describes how this project works.

The plotting module says to use plotly and not matplotlib. No `.py` source file
imports plotly — it appears only in notebooks and docs — and this file already
states the same rule in its own Plotting section, so the include was carrying a
duplicate of a rule the source tree does not exercise.

The pandas module is about cleaning survey data: build the frame column by
column, name each raw column at its assignment, reshape and merge only at the
end. pandas appears in 39 files here, but as DataFrame construction for library
output rather than as a cleaning pipeline; four files touch a CSV reader and
none misuse `inplace`. Its version premise did hold — the environment resolves
pandas 3.0.5 — so this is a relevance judgement, not a correction.

The jax module stays, and earns it: `custom_jvp`/`custom_vjp` occur 24 times in
`src`, `jax_enable_x64` 57 times across `src` and `tests`, and
`register_pytree_node` 7 times. Its one piece of unused advice is the argument
for keeping it rather than against — `lax.cond` and `lax.switch` appear zero
times, so the warning that `jnp.where` evaluates both branches and can
manufacture a NaN is guidance this code has not yet taken up.

Also advances the .ai-instructions pointer to pick up the codespell dictionary
and the pre-commit.ci skips for the submodule-backed hooks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
… bank

When a refined nested proposal cannot be replayed at a subject's realized
state, simulation falls back to the branches the solve published: the keeper
plus one conditional inner policy per outer mesh node. It now emits the
canonical-Q maximizer over the branches that are both reachable and
canonically feasible there, instead of the branch carrying the highest value
the solve happened to publish for it.

Three things change together, because each alone leaves a way for a valid
published branch to be discarded.

The selector returns the complete bank rather than one winner. Its rows are
node-major, keeper first, each carrying its own conditional inner action, and
its mask covers only replay structure, interpolation support, finiteness and
the solver-owned consumption budget.

`_nested_grid_baseline` then scores every surviving row through canonical Q in
one vmapped call and drops infeasible or non-finite rows before the argmax.
The winner used to be chosen first and scored afterwards, so a branch the
realized state makes infeasible could take the selection and then be rejected,
discarding the rest of the bank with it. Ranking on published values could
also prefer a branch that canonical Q ranks lower.

Both the bank predicate and the per-pair admissibility check now read one
domain, the published outer mesh. They used the declared inverse domain and
the mesh extrema respectively, so a keeper landing in the gap between a wide
declared domain and a narrower mesh was admitted by the first and rejected by
the second, suppressing a reachable node that nothing reconsidered.

Tie order is explicit: the keeper is row zero and wins an exact tie, then
ascending mesh-node order; the grid baseline wins an exact tie against the
best published row, preserving the no-degradation convention. When no branch
survives the caller raises before emitting an action.

The accompanying test module covers a lower-published-value feasible branch, a
published/canonical order reversal, both tie conventions, inner-policy
alignment, an all-infeasible bank, and a mesh narrower than the declared
domain, each at both precisions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Selection now ranks the keeper and every published node by canonical Q, so the
branch values the fixture published no longer decide which branch a fallback
emits. The module steered selection with those values, which left ten tests
asserting the superseded rule and three more passing whichever rule ran: the
stub objective `-|investment|` makes the keeper the unconditional argmax.

Each test now names the outer action its objective favours, and a refusal is
made observable by peaking the objective at the branch the rule must refuse --
without that the favoured branch loses on value anyway and the test reports the
same outcome whether the rule excluded it or not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq

@timmens timmens left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR 407 review comments

Reviewed against a5c6fe4fc.

Approving with one explicit merge condition: please fix the keeper-domain issue below.
Once that blocker is addressed and covered by a regression test, this is ready to merge
from my side without another review round.

Blocking: outer-policy replay correctness

Preserve the exact keeper outside a narrower adjuster mesh

AdaptiveOuterMesh solves the keeper as a separate exact branch, and the current
validation explicitly permits the adjuster mesh to lie strictly inside the outer
state's domain. Replay then applies the mesh endpoints to the keeper too. A keeper
outside that narrower adjuster mesh is removed before the canonical fallback ranks the
bank.

I reproduced this with the ordinary outer-state domain [0, 20], adaptive adjuster
mesh [2, 18], and a subject at outer stock 0. Terminal utility strictly prefers the
lower stock, so the solve chooses no adjustment and reports -8.230213. Replay rejects
that keeper solely because 0 lies below the adjuster mesh, sets
nested_policy_fallback=True, emits investment 2 and stock 2, and reports
-18.415411.

Could the keeper row be checked against the outer state's domain while adjuster rows
remain checked against the mesh? Requiring the adaptive mesh to span the full state
domain would also make the two phases agree, though preserving the separately solved
keeper seems closer to the current solver contract. An end-to-end regression should
cover a current stock outside a legal, narrower adjuster mesh.

Nonblocking: numerical robustness

Bound the round-trip error for interior outer targets

I reproduced an end-to-end solve/simulation disagreement on this head. Consider the
declared unit-slope map new_outer = outer + outer_investment, an outer state/search
grid (0.0, 0.5, 1e10), and a subject at outer=1e10. The solve ranks the interior
target 0.5 and reports a Bellman value of 0.693147. In float32, inversion recovers
the action -1e10, whose forward image is 0.0. The current predicate admits that
candidate because 0.0 lies somewhere inside the declared state domain. Simulation's
canonical rescoring consequently emits new_outer=0.0 and reports 0.455647.

This means the conditional problem is solved at one stock and replayed at a materially
different stock. Domain containment does not bound the interior approximation, and the
distance can be arbitrarily larger than the target's local ULP. Canonical rescoring
reports the value of the emitted action honestly, though it cannot repair the Bellman
value already computed at the nominal target during backward induction.

This requires an unusually wide and unevenly scaled state domain, so I do not consider
it merge-blocking. Could interior admission eventually enforce a target-local
round-trip bound, for example exact or adjacent-representable reproduction, and refuse
candidates outside that bound? An end-to-end regression with the values above would
pin that a published target and its replayed stock remain numerically associated.

@hmgaudecker
hmgaudecker merged commit b803c81 into main Aug 31, 2026
20 of 21 checks passed
@hmgaudecker
hmgaudecker deleted the feat/continuous-outer branch August 31, 2026 10:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants