Skip to content

NB-EGM: breakpoint-aware EGM solver for non-convex, discrete-continuous budgets - #400

Merged
hmgaudecker merged 1514 commits into
mainfrom
feat/nb-egm
Aug 26, 2026
Merged

hmgaudecker merged 1514 commits into
mainfrom
feat/nb-egm

Conversation

@hmgaudecker

@hmgaudecker hmgaudecker commented Jul 7, 2026 •

Copy link
Copy Markdown
Member

Summary

This PR has grown into the complete public and runtime surface for breakpoint-aware
endogenous-grid solving in pylcm. It adds the one-margin NBEGM solver and the nested
two-margin NNBEGM solver, but it also changes how models declare piecewise grids,
institutional boundaries, case-specific formulas, and solver-readable constraints. It
then carries that topology through continuation values, feasibility, simulation, native
envelope kernels, persistence, documentation, benchmarks, and CI.

The branch is based on the structural solver/constraint-routing work merged in #390.
Against current main, this PR changes 287 files (+40,474/-1,893).

User-facing API changes

Piecewise grids: breakpoint-first API (breaking)

PiecewiseGridSegment(interval=..., n_points=...) has been removed from the public API.
PiecewiseLinSpacedGrid and PiecewiseLogSpacedGrid now declare one closed domain plus
explicitly owned interior breakpoints:

from lcm import GridBreakpoint, PiecewiseLinSpacedGrid

grid = PiecewiseLinSpacedGrid(
    start=0.0,
    stop=100.0,
    breakpoints=(
        GridBreakpoint(value=10.0, owner="right"),
        GridBreakpoint(value=40.0, owner="left"),
    ),
    points_per_segment=(20, 50, 30),
)

This replaces interval-string parsing with a single canonical geometry:

  • outer endpoints are closed;
  • every interior breakpoint appears exactly once;
  • owner="left" / "right" says which adjacent segment contains equality;
  • the open side begins at the adjacent representable floating-point value;
  • each point count is the number of output nodes contributed by that nominal segment;
  • invalid counts and breakpoint layouts raise deterministic GridInitializationErrors.

The same ownership model is shared with NBEGM's feasibility and continuation topology,
including coincident left/right boundaries and strict constraints at an outer endpoint.
There is no compatibility shim for PiecewiseGridSegment; callers must migrate to
GridBreakpoint plus start, stop, breakpoints, and points_per_segment.

Structural declarations for non-smooth formulas

New decorators let model code retain the structure NBEGM needs without changing the
callable's behavior under GridSearch:

  • boundary(...) and case_boundary(...) declare a named comparison surface, its
    equality owner, and whether it is a continuous kink, jump, or hard constraint;
  • piece(...) attaches the smooth formula used on either side of a case predicate;
  • affine_breakpoint(...) and piecewise_affine(...) declare a piecewise-affine
    schedule, including indexed and MappingLeaf thresholds;
  • smooth_helper(...) distinguishes smooth numerical helpers from undeclared
    piecewise primitives.

The branch also exports fixed-form accounting declarations for the kernels it can prove
by identity: cash_on_hand_with_subsidy, liquid_law_from_savings, and
liquid_law_from_resources.

Constraints are planned, not implicitly discharged

NBEGM is integrated with the structural constraint route ledger. Every constraint must
be evaluated at a stage with all required inputs, proved by construction, compiled into
a boundary program, or rejected during model construction. Matching
post_decision_lower_bound declarations can be proved from the savings grid; suitable
Condition / ref(...) comparisons can be compiled into feasibility boundaries; an
opaque callable is accepted only where it is directly evaluable.

The NBEGM route keeps feasibility provenance and boundary ownership through the
within-period envelope, published continuation carry, nested NNBEGM solve, and simulated
policy. It covers intersections, one-sided constraints, coincident equalities in either
conjunction order, and strict lower/upper endpoint germs. Unsupported compositions fail
early instead of silently weakening the feasible set.

The new solver-selection documentation starts from the model author's question—when a
plain callable is enough and when retained Condition structure is needed—and then
explains which solvers can consume it.

Additional public surface

  • NBEGM and NNBEGM are exported from lcm.solvers.
  • SimulationResult.load_solution(directory=...) restores only saved value-function
    arrays, without loading the per-subject simulation checkpoint or metadata payload.
  • Phased now documents and enforces the perceived-law-versus-simulated-truth boundary,
    including phase-invariant constraints and restrictions on stochastic next-state reads.

Solver implementation

NBEGM

NBEGM solves a one-dimensional consumption/savings problem whose budget or feasible
set has declared institutional breakpoints. Its lowering and numerical paths cover:

  • case-piece budgets such as eligibility cliffs and notches;
  • piecewise-affine schedules with kinks, jumps, floors, indexed thresholds, and
    ride-along states;
  • discrete actions that enter utility, co-states, liquid laws, regime transitions, or
    schedule variables;
  • smooth and zero-breakpoint budgets;
  • secondary kinks from discrete choices;
  • stochastic continuation values and distributed state arrays.

Each smooth case is solved by EGM. Boundary-targeting and corner candidates are added
explicitly, then all case and discrete branches are combined with a query-side upper
envelope. jump_read="one_sided" publishes duplicated one-sided carry rows at value
jumps; "bridged" keeps an ordinary finite-grid continuation for faster approximate
reads.

NNBEGM

NNBEGM adds a finite outer keeper/adjuster search around an inner NBEGM solve. It is
the two-margin counterpart of NEGM: declared liquid kinks, jumps, and hard constraints
retain NBEGM treatment for every outer candidate, while the durable/illiquid margin is
searched on an exogenous grid plus the no-adjustment candidate.

Recursive preferences, arithmetic, and memory controls

  • NBEGM supports nonlinear Epstein–Zin recursion by applying the certainty equivalent
    to the joint regime-by-shock lottery, including composite single-power flow utility.
    The unsupported Epstein–Zin plus EV1 taste-shock combination is rejected explicitly.
  • Envelope ownership supports certified double-double comparison and a faster ordinary
    working-precision mode.
  • Fixed-width maps and separate batching controls for stochastic nodes, intervals,
    ride-along cells, discrete branches, candidate segments, and NNBEGM outer candidates
    bound intermediates without changing admitted results.
  • Parameter-dependent scope checks run once on first solve. Affinity and
    interval-constancy probes fail closed by default; probe_failure="assume_declared"
    is the explicit opt-out for models validated by another oracle.

Native kernels, examples, and operational tooling

The exact-affine CPU/CUDA kernels are now treated as an installed package payload rather
than an artifact left in the source tree. Ordinary and editable builds include a native
manifest fingerprinted by sources, Python ABI, platform, JAX/jaxlib, compiler toolchain,
CUDA toolchain, and architecture flags. Runtime/CI probes distinguish a missing or stale
payload (reinstall once) from an unloadable payload (fail).

This PR also:

  • adds a full NBEGM user guide and solver-selection guide;
  • rewrites the piecewise-grid documentation for the breakpoint-first API;
  • updates the Mahler–Yum example and benchmark to the current regime/grid surface;
  • pins and preflights the ACA benchmark dependency so API drift fails before ASV;
  • introduces declarative CI capability, coverage, resource, isolation, and tier markers;
  • bounds GPU PR coverage and shards the slow CPU solution suite in fp32 and fp64;
  • preserves JAX compilation caches and uploads policy/native/JUnit evidence.

Validation

The solver tests compare values, policies, marginals, discrete choices, and boundary
ownership against dense brute-force and independent envelope oracles. Coverage includes
case order, both equality owners, strict endpoint boundaries, carried feasibility,
one-sided and bridged jump reads, NNBEGM nesting, Epstein–Zin joint lotteries, batching
invariance, fp32/fp64 behavior, native-kernel capability, and distributed continuation
paths.

At exact branch tip 017fbcb1b5ee155615f662e953af63aa677d9ce7, the current GitHub
check set is green: Ubuntu fp32/fp64 (including all slow-solution shards), macOS,
Windows, GPU32, GPU64, ty, pre-commit, notebooks, documentation, benchmarks, Codecov,
and the ACA compatibility preflight.

@github-actions

github-actions Bot commented Jul 7, 2026 •

Copy link
Copy Markdown

Benchmark comparison (main → HEAD)

Comparing 88ce51dc (main) → abe354de (HEAD)

Benchmark Statistic before after Ratio Alert
aca-baseline execution time 10.208 s 9.991 s 0.98
peak GPU mem 604 MB 601 MB 1.00
compilation time 437.84 s 426.17 s 0.97
peak CPU mem 10.43 GB 10.55 GB 1.01
aca-baseline-debug execution time 36.832 s 35.450 s 0.96
peak GPU mem 602 MB 603 MB 1.00
compilation time 481.65 s 463.03 s 0.96
peak CPU mem 6.53 GB 6.36 GB 0.97
Mahler-Yum execution time 4.747 s 3.031 s 0.64
peak GPU mem 788 MB 685 MB 0.87
compilation time 15.97 s 21.75 s 1.36 ❌
peak CPU mem 1.35 GB 1.50 GB 1.11 ❌
Precautionary Savings - Solve execution time 26.2 ms 27.2 ms 1.04
peak GPU mem 34 MB 34 MB 1.00
compilation time 2.33 s 2.34 s 1.00
peak CPU mem 1.07 GB 1.07 GB 1.01
Precautionary Savings - Simulate execution time 62.0 ms 66.2 ms 1.07
peak GPU mem 328 MB 328 MB 1.00
compilation time 4.88 s 4.99 s 1.02
peak CPU mem 1.23 GB 1.23 GB 1.00
Precautionary Savings - Solve & Simulate execution time 100.1 ms 98.7 ms 0.99
peak GPU mem 1.68 GB 1.68 GB 1.00
compilation time 7.04 s 7.20 s 1.02
peak CPU mem 1.25 GB 1.25 GB 1.00
Precautionary Savings - Solve & Simulate (irreg) execution time 233.4 ms 231.0 ms 0.99
peak GPU mem 2.20 GB 2.20 GB 1.00
compilation time 7.51 s 7.55 s 1.01
peak CPU mem 1.32 GB 1.32 GB 1.00
IskhakovEtAl2017DCEGMSimulate execution time 236.9 ms 221.4 ms 0.93
compilation time 7.54 s 7.40 s 0.98
peak CPU mem 1.47 GB 1.46 GB 1.00
IskhakovEtAl2017DCEGMSolve execution time 1.906 s 1.886 s 0.99
compilation time 7.15 s 7.23 s 1.01
peak CPU mem 1.39 GB 1.37 GB 0.99
IskhakovEtAl2017Simulate execution time 225.2 ms 234.2 ms 1.04
compilation time 7.74 s 7.56 s 0.98
peak CPU mem 1.22 GB 1.21 GB 1.00
IskhakovEtAl2017Solve execution time 53.9 ms 52.2 ms 0.97
compilation time 2.63 s 2.62 s 1.00
peak CPU mem 1.11 GB 1.12 GB 1.01
IskhakovEtAl2017DCEGMSimulateGpuPeakMem peak GPU mem 440 MB 440 MB 1.00
IskhakovEtAl2017DCEGMSolveGpuPeakMem peak GPU mem 0 MB 0 MB 1.00
IskhakovEtAl2017SimulateGpuPeakMem peak GPU mem 440 MB 440 MB 1.00
IskhakovEtAl2017SolveGpuPeakMem peak GPU mem 67 MB 67 MB 1.00

@hmgaudecker

Copy link
Copy Markdown
Member Author

Round-4 external audit repairs (ledger: pr400-round4/REPAIRS.md in the audit archive):

  • F1 (serious) — the EZ kernels were discontinuous one ULP from gamma = 1 / rho = 1 (the log(sum)/(1-gamma) form lost ~47% at nextafter(1, inf)). The transform partials are now a quint (anchor, weight sum, anchored deviation sum, marginal log-scale, marginal mantissa); inversion is log nu = a + [log(W) + log1p(E/W)]/(1-gamma), exact and smooth through both unit limits. The auditor's suggested weight normalization was rejected: regime blends can carry genuinely sub-unit mass, so log(W) must be preserved. One-ULP tests at nextafter(1, ±inf) are permanent (0e44a2a8, b95b8727).
  • Full suite at b95b8727: 2030 passed, 56 skipped, 2 xfailed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01BNQkLG5QzXcgGdpgH6sN4u

hmgaudecker added a commit that referenced this pull request Jul 21, 2026
Brings the NestedNBEGM two-asset (Kaplan-Violante) solver and Epstein-Zin two-asset
support alongside the shared NBEGM machinery it drives (envelope_at_query asset-row
backend + nbegm_multi_interval_step_savings, already on this branch). Net-new files
are disjoint from the KV/MSS fixes: nnbegm.py, negm.py changes, certainty_equivalent,
user_regime_validation, solvers/regime wiring, plus n_nbegm / EZ / negm test suites.

Supersedes PR #403; the whole EGM surface is now reviewed under #400.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BNQkLG5QzXcgGdpgH6sN4u
hmgaudecker added a commit that referenced this pull request Jul 28, 2026
Everything that is not specifically NB-EGM/NNBEGM belongs on feat/dcegm.
These files carried #400 deltas that were never split across, so they are
taken wholesale from feat/dcegm rather than reconciled here:

- `egm/upper_envelope/mss.py` — #400's rounds-17-20 lineage retires in
  favour of feat/dcegm's version. The five #400-only MSS test files go
  with it; `test_mss.py` and `test_mss_fues_parity.py` revert.
- `egm/upper_envelope/{__init__,query}.py`, `egm/continuation.py`,
  `solution/dcegm.py` — reverted, so nothing here promises a
  read-support flag no backend produces.

What stays is NB-EGM content that happens to live in shared files: the
NBEGM/NNBEGM Epstein-Zin certainty-equivalent gate in
`user_regime_validation.py`, and the public API in `lcm/__init__.py`,
`solvers.py`, `exceptions.py`, `regime.py`, `case_piece.py`.

One known regression until the reverted `query.py` fix lands on
feat/dcegm and returns: `test_nbegm_savings_node_singleton_bracket.py::
test_envelope_brackets_a_lone_candidate_at_its_own_abscissa` reads
`env_value[0] == 1.0` where the contract is `5.0`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BNQkLG5QzXcgGdpgH6sN4u
hmgaudecker added a commit that referenced this pull request Jul 28, 2026
One conflict, in `regime_building/processing.py`, where two adjacent import
blocks collided:

* #405 adds `processes.ar1` / `processes.iid` / `processes.state_conditioned`
  imports for the state-conditioned sigma machinery;
* #400 carries the age-specialization -> age-normalization re-architecture, which
  moves the periodized API into `regime_building.age_normalization`
  (`AgeGridSchedule`, `Periodized*`, `resolve_periodized_nodes`,
  `periodized_tree_signature`, ...).

Resolution is the union of the two features, not of the two texts: keep #405's
`processes.*` imports and take #400's `age_normalization` block, dropping the old
`age_specialization` names (`_SpecializedEconFunction`, `resolve_state_grids`,
`tree_signature`, ...). Those are genuinely gone from this module's body — the
merged code calls `periodized_tree_signature` / `resolve_periodized_nodes`
throughout, and no bare reference to a dropped name survives. `age_specialization`
itself is NOT deleted upstream; it remains the lower-level helper module that
`age_normalization` and `model_processing` build on, so its other importers are
unaffected.

Checked both payloads survive the merge, since earlier cascades of this same pair
dropped bindings: all 208 lines #405 added to this file and all 287 lines #400
added are present afterwards, and every name imported by the resolved block is
used in the body.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DW9D1WhLN8JhhiEPCScbaw
hmgaudecker added a commit that referenced this pull request Jul 29, 2026
`Run tests on GPU (32-bit precision)` has been red on this PR with six failures in
`test_state_conditioned_shocks.py`. All six are the tests pinning x64-grade
tolerances, not a defect in the rows:

* four `assert_allclose(..., atol=1e-10)` comparisons of our row against pylcm's
  OWN CDF-binned row. The gap is floating-point noise in shared arithmetic, and
  under float32 it is a few ULPs on O(1) probabilities — measured max |diff|
  ~3e-7 against eps = 1.2e-7. `atol` is now `1e-10` under x64 and `1e-6`
  (~8 ULP) under float32, which is still a tight bound, not a blanket loosening.
* `float(gather_sigma(...)) == 0.2`, which cannot hold in float32 because 0.2 is
  not representable there (it round-trips to 0.20000000298023224) — now
  `pytest.approx`.

Precision comes from `tests.conftest.X64_ENABLED`, the same switch
`test_processes.py` already uses, so it follows `--precision`.

Reproduced the six failures locally under `--precision 32` before touching
anything, so the fix is verified against the actual failure rather than assumed:
now 48 passed under `--precision 32` AND 48 passed under the x64 default. ruff
(pinned v0.15.20), ruff-format and ty are clean.

Note the two remaining GPU-32bit failures are NOT from this branch:
`tests/egm/test_crra_utility.py` and
`tests/solution/test_nbegm_nonpositive_corner.py` arrive with feat/nb-egm (#400).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DW9D1WhLN8JhhiEPCScbaw
hmgaudecker added a commit that referenced this pull request Jul 29, 2026
Taken from `feat/dcegm` (PR #390, which is where this change actually lives --
PR #400's `pyproject.toml` is byte-identical to this branch's, so cherry-picking
from #400 would have been a silent no-op).

ONLY the aarch64 hunks are taken. #390's `pyproject.toml` diff also reorders the
ruff `per-file-ignores` table and drops the `ARG001` ignores for
`tests/test_certainty_equivalent.py` and `tests/test_temporal_aggregation.py`;
both files exist here and use those fixtures, so importing that churn would red
the lint for reasons unrelated to CUDA.

`pixi.lock` is regenerated rather than patched, because adding a platform forces
a re-solve. Checked, not assumed:

- the aarch64 wheels are the right interpreter --
  `jax_cuda13_plugin-0.11.0-cp314-cp314-manylinux_2_27_aarch64.whl` against this
  workspace's py3.14;
- `nvidia_cudnn_cu13-9.24.0.43-...aarch64.whl` resolves. This is the one that
  matters: without cuDNN, jax loads the plugin, logs a line, and falls back to
  CPU roughly 20x slower -- a "GPU run" that silently is not one. Any run on that
  box should still assert `jax.default_backend() == "gpu"`;
- no existing platform lost packages -- linux-64 2621 -> 2621, win-64 615 -> 615,
  osx-arm64 592 -> 592. The deleted lock lines are block reordering from
  inserting linux-aarch64, not removals;
- `tests/solution/test_envelope_query.py` still passes locally under the
  regenerated lock (10 passed on the round-9 scale regressions).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Jul 30, 2026
The `feat/nb-egm` -> `feat/state-conditioned-shocks` merge (2ba4237) resolved the
`broadcast.py` conflict correctly where it was visible — nb-egm's `ages=ages`
threading plus this branch's `_state_conditioned_names` block, both kept — but
silently lost 21 lines of nb-egm's payload in the same two files:

* `broadcast.py`: the `resolve_node` import AND the whole
  `_resolved_at_representative_age` helper, while keeping its two call sites, so
  `_needed_names` raised `NameError` on every `Model(...)` that prunes broadcast
  variables;
* `model_processing.py`: `from _lcm.regime_building.age_normalization import
  _regime_has_markers`, still called at :330.

Both are the same mechanism: the two sides inserted an import at the same sorted
position, and the resolution took one line instead of both. Nine tests were red
across `test_model_broadcast.py`, `test_state_conditioned_model.py` and
`regime_building/test_age_normalization.py`; all pass now.

Restored verbatim from `feat/nb-egm` — this is nb-egm's code, unmodified by this
branch, so there is nothing to re-derive.

Verified by line-presence against the merge-base of the two parents (NOT against
`HEAD...origin/feat/nb-egm`, which is empty once HEAD contains nb-egm and so
proves nothing): #400 10759 added lines, 0 missing; #405 1660 added lines, 0
missing. Tests: the three files this PR adds, the four it modifies, and everything
exercising the merge-resolved modules (broadcast / model_processing /
regime_building, ages, phased, stochastic, params, age-specialized solves) —
371 passed; the PR's own tests also 48 passed under `--precision 32`. ruff
(pinned v0.15.20), ruff-format and ty clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DW9D1WhLN8JhhiEPCScbaw
@codecov

codecov Bot commented Jul 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.15315% with 114 lines in your changes missing coverage. Please review.
✅ Project coverage is 92.61%. Comparing base (88ce51d) to head (abe354d).

Files with missing lines Patch % Lines
hatch_build.py 51.66% 29 Missing ⚠️
src/_lcm/simulation/simulate.py 89.51% 13 Missing ⚠️
src/_lcm/solution/nnbegm.py 95.20% 12 Missing ⚠️
src/_lcm/egm/nbegm_validation.py 87.64% 11 Missing ⚠️
src/_lcm/egm/upper_envelope/_exact_affine/ffi.py 82.45% 10 Missing ⚠️
benchmarks/preflight.py 0.00% 5 Missing ⚠️
src/lcm_examples/mahler_yum_2024/__init__.py 82.14% 5 Missing ⚠️
src/_lcm/egm/continuation.py 69.23% 4 Missing ⚠️
src/_lcm/grids/piecewise.py 96.49% 4 Missing ⚠️
src/_lcm/regime_building/finalize.py 92.72% 4 Missing ⚠️
... and 10 more
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #400      +/-   ##
==========================================
+ Coverage   91.43%   92.61%   +1.18%     
==========================================
  Files         249      279      +30     
  Lines       24630    29135    +4505     
==========================================
+ Hits        22520    26983    +4463     
- Misses       2110     2152      +42     
Flag Coverage Δ
cpu-python 92.61% <93.15%> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

hmgaudecker added a commit that referenced this pull request Jul 31, 2026
Propagates the age-specialized/EGM continuation-grid cascade fix into
feat/continuous-outer. Five files conflicted; #400 and #407 have independently
rewritten the query-side envelope certification with no shared symbols beyond
the entry point and `_SegmentLinks`.

Per the resolution recorded in
/tmp/pylcm-fleet-coordination/envelope-architecture-390-400-407.md, #407's
exact-dyadic kernel is the implementation and is to be mounted at or below #400,
so the architecture side of each conflict resolves to ours:

- src/_lcm/egm/upper_envelope/query.py   (--ours)
- src/_lcm/egm/nbegm_step.py             (--ours)
- tests/solution/test_envelope_query.py  (--ours)

The remaining two conflicted files are NOT architecture. `git log
4a53546..114b69b -- <file>` shows both carry only 9fae051, "Read the NB-EGM
continuation on each period's own grid" -- the cascade fix itself. Its change
there is additive fixture plumbing: optional `illiquid_grid` and `liquid_grid`
kwargs that the fix's own regression test passes. That test,
tests/solution/test_age_specialized_solver_composition.py, is not among the
conflicted files, so it lands here and calls both kwargs; under `--ours` it
would raise TypeError and the fix would arrive with its regression dead.

Both parents' edits to those two files are disjoint -- #406 adds the grid
kwargs, #407 adds outer_search/branch_aggregator/dead_functions -- so they
resolve as a union that drops nothing from either side:

- tests/test_models/n_nbegm_toy.py       (union)
- tests/test_models/nbegm_common.py      (union)

Verification on the merge result. Hooks do not re-run on a merge commit, so
`prek run --all-files` was run by hand: it caught 2x ruff FURB164 in
test_envelope_query.py -- #407 code newly linted under #406's ruff config, red
in neither parent -- fixed here. `ty` clean. `pixi install --locked` clean for
default and tests-cpu; `pixi lock` a no-op. Every line added by either parent is
present in the merge across all 14 both-touched files; the 111 theirs-only and
57 ours-only changed files are byte-identical to their source parent.

test_age_specialized_solver_composition.py 9 passed. The seven #400
certified-sign/double-double module test files 76 passed, 2 skipped;
certified_sign.py, double_double.py and cell_hull.py land intact and do not
depend on query.py, so both architectures coexist without a dangling import.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Brings across #400's `06bf038` (pin the envelope owner across eager, JIT and
vmap) and `a60a092` (bound a certified margin by its own residual), #390's
`af4b05b`/`ec96c12`/`ffb25e2`/`cdd4f65`, and the merges that carried them.

One conflict, and it is NOT architecture: `_build_Q_and_F_per_period` in
`src/_lcm/regime_building/processing.py` gained a keyword-only parameter on both
sides -- `reachable_targets` from #390's declared-reachability work, and
`period_to_regime_v_interp` from this branch. Both are needed and both are
keyword-only, so they union; nothing is dropped. Per-file derivation with
`git log 114b69b..18f5491 -- <file>` confirms no other conflicted file, and in
particular `query.py` did not re-conflict: the ownership resolution from the
previous merge holds.

`tests/solution/test_envelope_margin_bound_is_outward.py` is REMOVED here, and
this is the one judgement call in the merge. It arrives with `a60a092` and
imports `_value_quotient` from #400's `query.py` -- an internal of the envelope
architecture that the recorded three-branch resolution supersedes on this
branch, so the symbol does not exist and `ty` fails. The property it asserts,
that a certified margin's bound is outward and never understates its own error,
is NOT superseded and applies to this kernel too; `certified_quotient_margin`
is present here. Porting it against this branch's `_dd_div` is owed work, filed
in the fleet coordination directory rather than dropped.

Verification: no line added by either parent is absent from the result (the one
both-touched file checked line-by-line; 11 theirs-only and 66 ours-only files
byte-identical to their source parent); `prek run --all-files` clean, run by
hand because hooks do not re-run on a merge commit; `pixi install --locked`
clean for default and tests-cpu; `pixi lock` a no-op.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Ten commits in one hop, held deliberately until #400's seven had reached #406
so this would not be two sequenced merges: `bf61b145` and its chain, `8f31991`
(publish a realized target reading a phase-invariant next state), `5b4339e`,
`d2f7b56`, `0b6be03`, `d9aec7d`, `3934701`, `f7ae72a`.

ZERO conflicts, which is where silent drops hide, so the per-file rule was
applied to the result rather than skipped: 16 theirs-only and 66 ours-only
changed files are byte-identical to their source parent, and all three
both-touched files (`regime_building/processing.py`, `solution/nbegm.py`,
`tests/simulation/test_simulate.py`) were checked line-by-line -- every line
added by either parent is present in the merge.

`model_processing.py` keeps the `from _lcm.processes.base import
_ContinuousStochasticProcess` form that #400->#405 settled on as strictly
narrower and cycle-safe; it did not resurface as a conflict here, and the module
was confirmed by importing it, not by reading it.

Verification: `prek run --all-files` clean, by hand, because hooks do not re-run
on a merge commit; `pixi install --locked` clean for default and tests-cpu;
`pixi lock` a no-op; 137 passed across the certificate guard, the cascade
regressions and `tests/regime_building`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 1, 2026
Brings the combined cascade hop to #409: bf61b14 (#400), 8f31991 (#406
transition-dependence fix), ddbec5b and 5b4339e are now all ancestors.

One conflict, in `Q_and_F.py`, and both sides were right about different things.

#407 `d9aec7d` transforms a stateless target's value BEFORE the expectation:
`QuasiArithmeticMean` is `CE = g^-1(sum_r p_r * E_w[g(V'_r)])`, so `transform`
precedes every expectation including the regime-transition one, and `inverse`
applies once to the finished sum. Previously a stateless target was
probability-weighted raw and only then inverted, leaving the transformed domain.

#409 routes stateless targets through `mixture_terms` rather than a separate
`E += p*V`, so they get `zero_safe_weighted_term` (a zero-probability +-inf is
masked to exactly 0 instead of yielding NaN) and the value-ordered,
alpha-renaming-invariant reduction. On this branch dissolution makes `-inf`
continuations ordinary, so a raw product is a live NaN source, not a hypothetical.

Resolved by doing both: CE-transform the stateless value, then seed
`mixture_terms` with the transformed value. `_sum_regime_mixture` forms
`sum_r p_r * g(V_r)` zero-safely and `ce.inverse` is applied once to the result,
which is exactly `g^-1(sum_r p_r * g(V'_r))`.

Also drops `test_collective_solve_default_call_still_returns_bare_mapping`
(external test-suite dedup audit, round 1): it is a local mirror of
`test_collective_regime_simulate.py::test_public_model_solve_default_return_shape_is_byte_identical`,
which is the module it already imports its model helper from, and whose
assertions are a strict superset. The file's own singleton guard,
`test_singleton_solve_bare_call_returns_legacy_mapping_not_a_tuple`, is retained
-- that is the invariant this module exists for. The other five deletions from
that audit are shared with `main` and went to #390 as a patch.

Verified: ty clean; tests/regime_building + the two touched test files 379 passed;
the certainty-equivalent bearing files (test_temporal_aggregation,
test_dcegm_validation, test_nbegm_epstein_zin_composite_flow,
test_simulate_terminal_rows) 43 passed.

KNOWN COVERAGE GAP, not closed here: no test appears to construct a model with
BOTH a certainty equivalent AND a stateless target, so the intersection this
resolution creates is exercised by neither side's tests. `d9aec7d` itself shipped
without tests. The two halves are covered separately; their combination is not.
hmgaudecker added a commit that referenced this pull request Aug 4, 2026
…ansitions.

This carries nb-egm (#400) through #405. Three conflicts, none resolvable by
picking a side: the incoming branch REPLACED mechanisms this branch had
extended, so taking either parent whole would have dropped the other's fix.

`Q_and_F.py` -- the incoming branch extracted the inline continuation loop into
`_get_compute_E_next_V(functions=...)`. This branch's whole feature is that
those two builds read `continuation_pool`, NOT `functions` (the continuation is
priced under the perceived solve-phase law). Adopted the extraction and passed
`functions=continuation_pool`. Verified `functions` is used at exactly the two
continuation sites inside the callee and nowhere else, so the substitution is
exact; in the solve phase the two are the same object. Recorded the contract in
the callee's docstring so a later flow-pool use cannot be added silently.

`processing.py` -- five of six hunks were UNIONS, not either/or: the incoming
`phase_reachability`/`source_regime_name` and this branch's
`coarse_state_law_names` are different parameters serving different concerns,
and `coarse_state_law_names` is still read in the auto-merged body. The sixth
replaced the inline `continuation_info`/`group_key` with
`continuation_info_lookup`/`continuation_group_key`, neither of which knew about
this branch's per-period perceived pool. Adopted the helpers and extended
`continuation_group_key` with `continuation_functions`, folding its per-period
signature into the key (round-11 F1: periods whose perceived pool resolves to
different closures must not share a compiled kernel). `None` contributes a
constant, so the diagnostics caller -- which never folded it -- is unchanged.

`tests/test_Q_and_F.py` -- import union.

One further break surfaced only in the test run: nb-egm's new guard
`test_active_periods_are_computed_at_a_single_call_site` fired, because this
branch's round-12 F1 constraint-ancestry check computed
`ages.get_periods_where(user_regime.active)` itself. That is exactly the second
evaluation point the guard forbids. Threaded the prepared
`active_periods_by_regime` into `_validate_constraint_phase_invariance` and read
it there, mirroring `_validate_all_variables_used`; the sole production caller
always supplies it.

Drop check on the conflicted files: 0 of 324/476/261 incoming added lines
missing. 15 of this branch's lines are absent by exact match, all of them the
two blocks deliberately relocated above (the `continuation_pool` calls and the
`continuation_functions_sig` block) -- enumerated and accounted for, none lost.

Verified: prek --all-files green (ruff, ruff-format, ty); 357 tests across
test_Q_and_F, test_coarse_law_provenance, tests/regime_building/,
test_certainty_equivalent, simulation/test_reachability_phase_split and
solution/test_reachability; plus all of tests/simulation/ (131 passed, 2
skipped, 5 xfailed), which is where `continuation_pool` and `functions` actually
differ. Not the full suite.
hmgaudecker added a commit that referenced this pull request Aug 4, 2026
Carries nb-egm (#400) and state-conditioned shocks (#405) through #406. Four
conflicts.

`src/lcm/solvers.py` and `src/lcm/__init__.py` are a RENAME collision, not an
export union: nb-egm renamed `one_asset_egm.py`/`OneAssetEGM` to
`egm.py`/`EGM` and `two_dim_egm.py`/`TwoDimEGM` to
`two_asset_egm.py`/`TwoAssetEGM`, which this branch predates. Git applied both
renames and left this branch's stale import names behind, so a naive union
would have re-exported two names for classes that no longer exist. Adopted the
new names and kept this branch's own additions (`OuterSearch`,
`OuterBranchAggregator`, `LegacyGoldenSection`, `UniformObservedFixedCost`,
`AdaptiveOuterMesh`, `FiniteOuterGrid`, `DeterministicOuterMaximum`,
`BranchAggregateResult`). Verified this branch made NO change to either renamed
file between the merge base and its tip, so the rename costs it nothing.
`_lcm/egm/one_asset_egm_step` is a different module and is untouched.

`src/_lcm/solution/contract.py` -- import union, both symbols used.

`src/_lcm/regime_building/processing.py` -- `_build_Q_and_F_per_period` kept
this branch's `period_to_regime_v_interp` (still read at three sites in the
merged body) and dropped `reachable_targets`, which the incoming
`phase_reachability` refactor replaced: it survived only as its own parameter
declaration, and neither of the two call sites passes it.

`src/_lcm/egm/upper_envelope/query.py` is unchanged, blob `e990fb68` -- the
object under Pro audit, per the settled ownership decision that this branch's
exact-dyadic kernel is the implementation.

Drop check on all four conflicted files: 0 missing added lines on either side.

Verified: prek --all-files green (ruff, ruff-format, ty); 242 tests across
test_solvers, test_outer_search, egm/test_branch_aggregation,
test_continuous_outer_audit_regressions, tests/regime_building/, test_Q_and_F
and solution/test_reachability. Not the full suite -- CI covers that.
hmgaudecker added a commit that referenced this pull request Aug 4, 2026
Carries nb-egm (#400), state-conditioned shocks (#405) and perceived stochastic
transitions (#406) through #407. Two conflicted files, 15 hunks, and the
resolution is a PORT rather than a side-pick: the incoming branches refactored
exactly the machinery this branch had extended.

KNOWN RED: 13 tests still fail, all one class, see the last section.

## Q_and_F.py (5 hunks)

The incoming extraction `_get_compute_E_next_V` is adopted at BOTH call sites
and this branch's two inline aggregation blocks are gone -- but their numerics
were ported INTO the helper, not dropped:

- Linear path: `zero_safe_average` within a target, and the cross-target mixture
  reduced ONCE by `_sum_regime_mixture` -- stack the operands, one zero-safe
  contraction, value-ordered sum. That preserves round-8 (accuracy: stacking the
  operands rather than the formed products lands on the exact-policy side of the
  pinned 5-target fixture) and round-10 F1 (the reduction order is a function of
  the contribution multiset, never of regime LABELS, so an alpha-renaming cannot
  move the result). Upstream accumulated `E + p*V` in label order, which would
  have reintroduced both defects -- and `0 * -inf = nan` besides, which is
  ordinary here because dissolution makes `-inf` continuations routine.
- `_scalar_target_contribution` now returns UNMULTIPLIED terms, so a stateless
  target joins that same single contraction instead of being summed separately.
- CE path keeps UPSTREAM's single joint `ce.aggregate`. That deliberately
  supersedes this branch's per-target `transform -> reduce -> inverse`, which
  cannot anchor `PowerMean`; flattening before aggregating is the whole point of
  the incoming fix.

`get_period_targets` and `get_period_scalar_targets` are DELETED. They had no
live callers left after the merge, and #400 ships guards that forbid them by
name (`test_engine_has_no_period_target_inference_helper`) and ban the literal
`carry_targets` (`test_continuation_targets_are_not_derived_from_law_bundle_keys`).
Both guards pass. Comments that named them were reworded rather than left to
resurrect the banned vocabulary.

Git had also FUSED two valid states into a broken one here: it aligned
`def _get_compute_E_next_V(` with this branch's `get_period_targets` tail on the
two lines they share, splicing one function's body inside the other's signature.
Reconstructed deliberately rather than patched in place.

## processing.py (10 hunks)

Mostly unions -- the incoming `phase_reachability`/`source_regime_name` and this
branch's `coarse_state_law_names`/`fold_only_regimes` are different parameters
serving different concerns, and both are read in the merged body. Substantive
ones:

- `reachable_targets` (this branch, still read by the fold-only empty-bundle
  rule) and `continuation_targets` (incoming, the graph-retained process
  scoping) are BOTH kept: different sets, both live. The coarse-candidate
  `_fail_if_coarse_candidate_folds_ambiguously` guard is preserved.
- The collective builder takes `period_targets=stateful_targets` from
  `partition_continuation_targets`, which is exactly what this branch's
  `period_targets` has always meant.
- Removed a duplicate `source_regime_name` the union introduced at two call
  sites and in `_process_regime_core`'s signature (caught by `ty` as a genuine
  `invalid-syntax` duplicate keyword).

## Tests

38 `process_regimes` call sites migrated to #400's now-required
`prepared_structure`, via `tests.conftest.build_prepared_structure`, including
the `_solve_kwargs` helpers in `test_fold_gate_guard.py` and
`test_fold_guard_complete.py`.

## KNOWN RED: 13 failures, one class, remedy already decided

#400's new build-time handoff check (`2bfc333`, adopted by `94e868e`) rejects a
target that declares a stochastic process the source neither carries nor
supplies an entry law for. This branch's gated-edge / fold models exercise
exactly that shape -- the deleted `get_period_targets` docstring named it a
"known remaining hole (fold-review F2, deliberately NOT closed in this slice)",
and #400 closed it by rejecting such models.

Affected: `test_fold_iid_shocks.py`, `test_gated_edge_arg_provenance.py`,
`test_fold_guard_complete.py`, `test_fold_gate_guard.py` (31 instances of the
same rejection).

Decided remedy (maintainer): fix the MODELS, not the check -- give each source a
target-specific ENTRY LAW, `state_transitions={"wage_shock": {"target": <law>}}`
(shape already used at `test_fold_iid_shocks.py:471`). An entry law leaves the
source's state space, and so the fold/gate topology under test, unchanged.
Applied in the follow-up rather than here, because `feat/nb-egm` has moved
(`4cb10b6` -> `fdd5d6d`) and this merge is redone against the new tip anyway.

Verified: `prek run --all-files` green (ruff, ruff-format, ty); 597 of 612 pass
across tests/regime_building/, test_Q_and_F, test_collective_regimes_draft and
test_certainty_equivalent. Not the full suite. NOT pushed: the branch is
knowingly red on the class above.
hmgaudecker added a commit that referenced this pull request Aug 5, 2026
Brings in the aggregator specification objects, the comment-only pass, and
the EGM preference-map refactor (`feat/nb-egm` @ `48d3795c`, carrying
`feat/dcegm` @ `ae23e2f0`).

One textual conflict, and it is the class #400 warned about: an import block
where this branch's `_lcm.processes.state_conditioned` import sits adjacent
to upstream's widened `_lcm.transition_laws` import. `-X ours` would have
dropped upstream's `is_interpolation_basis`. Resolved by keeping both, and
verified by use rather than by assumption -- `is_interpolation_basis` is
read at `next_state.py:190`, `StateConditioned` / `gather_sigma` /
`sigma_array_by_code` at lines 400-505.

Checked the three clean-merge semantic breaks #400 flagged:

- No surviving identity check against the deleted `W_linear` /
  `W_epstein_zin`. Every `koopmans_aggregator is ...` hit is a legitimate
  `None` test or an identity assert against a user-supplied object. The
  NBEGM/NNBEGM gate arrived as upstream wrote it,
  `isinstance(regime.koopmans_aggregator, CESAggregator)`.
- `_lcm/egm/crra.py` stays deleted; `nbegm_step.py` reads `crra_utility`
  from `_lcm/egm/nbegm_crra.py`. No second copy added.
- Removes the stray 4.7 KB root artifact `hvgaudecker@snellius.surf.nl`
  that `da08033f` introduced, so it does not propagate to the three
  branches below this one.

Verification is the decisive set from #400 §5 plus this branch's own two
modules: 59 passed, none on the `RESTRICTION-cheap-tests-only.md` AVOID
list, asserted to import `_lcm` from this worktree's own `src`.
`prek run --all-files` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 5, 2026
…ansitions.

Carries #405's merge of `feat/nb-egm` @ `48d3795c` (aggregator specification
objects, EGM preference-map refactor, the runtime inactive-target check).

Two textual conflicts, both the disjoint-import class #400 flagged:

- `Q_and_F.py` -- this branch's `all_stochastic_next_state_names` against
  #405's `SupportAxes` / `is_interpolation_basis`.
- `tests/test_Q_and_F.py` -- this branch's `_law_sources_differ` against
  #405's `_regime_mass_is_unit` / `_unit_regime_mass_or_nan`.

Resolved as the union in both cases, and each symbol checked twice over: it
occurs in the merged body beyond its own import line, and its definition
resolves (`Q_and_F.py:1077/1431/1475`, `transition_laws.py:146`). `-X ours`
would have dropped the incoming three and `-X theirs` this branch's two;
neither is right, which is why the rule is to derive per file.

`ISC004` fires on two of this branch's own error strings in
`user_regime_validation.py` under the ruff config that arrived with the
merge. Wrapped both in explicit parentheses rather than silencing the rule:
these are entries in a returned `list[str]`, exactly the position where a
dropped comma would silently fuse two distinct validation messages into one.

Verified with the decisive set from #400 §5 plus this branch's own modules:
133 passed, none on the AVOID list, `_lcm` asserted to resolve to this
worktree on the controller and both xdist workers.
`prek run --all-files` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker added a commit that referenced this pull request Aug 5, 2026
Carries #405 and #406, and through them `feat/nb-egm` @ `48d3795c`.

Two textual conflicts:

- `src/lcm/__init__.py` -- the exact "keep your addition, take their
  deletion" case #400 warned about. This branch added
  `BranchAggregateResult` and `UniformObservedFixedCost` to `__all__` while
  upstream removed `W_epstein_zin` and `W_linear`. `-X ours` would have
  re-exported two names that no longer exist. Verified by importing the
  package and checking every entry resolves: 59 exports, no missing
  attributes.
- `src/_lcm/egm/nbegm_step.py` -- an import block. Took upstream's
  `crra_utility` from `_lcm/egm/nbegm_crra` (used 9x here; `_lcm/egm/crra.py`
  stays deleted, no second copy added) but NOT `affords_an_action`.

That second exclusion is deliberate and worth recording, because the naive
union looked right. `affords_an_action` is used at five sites on #406 and at
none here, which has the shape of a merge silently dropping a change. It is
not one: `db9cbdc0` is an ancestor of *both* branches, so it sits in the
merge base, and this branch later re-inlined the predicate on its own side
while #406 did not touch those lines -- git took the newer side correctly.
The two forms are also identical by construction, `affords_an_action` being
`return budget > 0.0`. So there is nothing to restore, and importing a
symbol this branch does not use would only have tripped `F401`.

Confirmed the three aggregator gates stay distinct rather than being
unified, as #400 §2 requires: `preferences.py:214` asks
`isinstance(solve_W, LinearAggregator)` for the EGM route,
`user_regime_validation.py:553` asks `isinstance(..., CESAggregator)` for
the NBEGM/NNBEGM Euler inversion, and `nbegm.py:4660`
`_aggregates_nonlinearly` keys off the certainty equivalent and excludes
`LinearExpectation` explicitly. No surviving `W_linear` / `W_epstein_zin`
reference anywhere in `src/` or `tests/`.

Verified with #400 §5's decisive set plus this branch's cheap kernel tests:
57 passed, none on the AVOID list, `_lcm` asserted to this worktree on the
controller and both workers. The five AVOID-listed continuous-outer modules
(21s to 9154s) are left to CI. `prek run --all-files` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
hmgaudecker and others added 6 commits August 14, 2026 05:34
Dekker two_prod is exact only while both splitting halves are numbers the
format can multiply. An operand within a significand width of the smallest
normal has the lower half of its own split below that threshold, and a
subnormal half reads as zero to the very multiplications the error term is
assembled from — while its true contribution against a large second
operand is an ordinary number. The transform then returns a decomposition
that is not the product and reports nothing, so every bound downstream
bounds the wrong quantity and a determinant arrives claiming an exactness
it never had.

On the witness operands the discrepancy is exactly the term that went
missing: `nextafter(tiny)` against the largest normal loses 2**-51.

`_split` already borrows range at the top for the mirror reason — the
splitting constant overflows before the product does. This is the same
borrowing at the bottom: the operands are moved apart before splitting,
one up by as much as the other goes down, so the product is the same real
number and only the halves change. Where both operands are that small the
shifts cancel, correctly, since the product is then out of range anyway
and its caller bounds it. Both scalings are round-trip checked and the
pair keeps its original operands where either is inexact.

Against the exact rational sign over the same 21600 geometries: 16198
strict verdicts before and after, with the four inverted signs gone and
the eighteen fabricated ties untouched. The strict count not moving is
what separates this from a repair that stops deciding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
`_subnormal_operand_present` refused every row carrying a subnormal
operand unconditionally, on the belief that XLA always flushes. CUDA
reads the whole band, so on that backend the refusal withheld a verdict
the arithmetic had reached correctly and published NaN in place of a
stored value the format holds exactly.

It now consults `backend_flushes_subnormals` and leaves the compiled
program where the backend reads the band, the way
`_derived_subnormal_possible` already does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
Verbosity flags reshape pytest's human-readable output — `-q` suppresses
the collection header entirely — so a grep over it reports on the format
rather than on the run. Every pytest invocation now also writes a JUnit
report, and each is uploaded as an artifact under `always()` so a failing
run still publishes the counts that explain it.

Paths are unique per invocation rather than a single fixed name: a run
that dies early otherwise leaves the previous run's report in place,
which reads as a valid result for a run that never finished.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
0.3.11 is the release in which jaxtyping's import hook handles
`@no_type_check` correctly. The cloudpickle sentinel fix this project
relies on landed one release earlier, in 0.3.10, so the two floors
answer different questions and only the later one covers both.

The lock moves jaxtyping in all seven environments and no other
artifact or version line with it; the remaining churn is purl source
annotations pixi rewrote on the same run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
hmgaudecker added a commit that referenced this pull request Aug 23, 2026
@hmgaudecker
hmgaudecker dismissed a stale review August 24, 2026 05:47

???

@timmens timmens left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Happy with all the changes. I only found three remaining small blockers

Blocking: Use the declared lower savings endpoint in the interval route

Placement: src/_lcm/egm/nbegm_step.py, _interval_corner_candidates

_interval_corner_candidates constructs its lower corner as s = 0: consumption is
corner_coh_grid, the continuation is read from interval_value[0], and the resulting
policy therefore implies zero savings. The public contract accepts a positive savings
lower bound and checks post_decision_lower_bound(...) against savings_grid[0].

A public ride-along model with a declared lower bound of 1 and a savings grid starting
at 1 passes model construction and log_level="debug". The interval route can then
publish an infeasible zero-savings choice; its solved value differs from GridSearch by
up to 2.9586 in the witness.

Could this corner use savings_grid[0] consistently for consumption, value, policy,
and feasibility, with a regression through the public interval-specific route?

Blocking: Probe affinity and unit slope inside every resolved interval

Placement: src/_lcm/solution/nbegm.py, _liquid_affinity_samples and
_liquid_slopes

The budget preflight does not verify every interval that the kernel reconstructs.
Second derivatives are sampled at liquid-grid nodes and their midpoints, without using
the resolved economic breakpoints. The all-jump unit-slope check ignores those samples
and evaluates the slope only at liquid values 1 and 2.

An all-jump schedule whose cash-on-hand slope is 1 below a threshold at 10 and 2
above it passes the check. NBEGM reconstructs the upper branch with unit slope, and the
accepted two-period model differs from GridSearch by up to 0.8676 under
log_level="debug".

Could the parameter check resolve each branch's breakpoint partition for the current
draw and evaluate both curvature and required slope at an interior point of every
nonempty interval?

Blocking: Replay the outer candidate selected by NNBEGM

Placement: src/_lcm/solution/nnbegm.py, _NNBEGMPeriodKernel.__call__

NNBEGM solves the outer decision over NNBEGM.outer_grid, then returns no simulation
policy. Simulation performs a fresh grid search over the regime's outer action grid,
which can represent a different candidate set from the one used during solution.

Changing only that action grid leaves the supplied solved value unchanged while
changing the simulated result:

                         narrow action grid   wide action grid
Solved value                 -7.8459              -7.8459
Simulated value              -9.9077              -7.8841
Chosen outer action            0.01                 5.00

Could NNBEGM publish and replay the joint outer and inner policy, or validate during
model construction that the simulation action grid represents exactly the solver's
candidate set?

@hmgaudecker

Copy link
Copy Markdown
Member Author

Thanks — all three blockers are addressed:

  • The interval corner now uses the declared savings-grid lower endpoint consistently, with public ride-along/cliff regressions (78dfa1a).
  • Budget validation now resolves the economic breakpoint partition and checks curvature plus required unit slope inside every live interval, including the threshold-at-10 witness in fp32/fp64 (12ce852).
  • NNBEGM now publishes its aligned keeper-plus-outer joint candidate bank and simulation replays the solved winner independently of the simulation action grid; the scalar oracle, batching/refinement checks, and runtime/memory gates are green (3758232).

The exact-tip blocker witnesses pass 7/7; all-file ty and commit hooks are green.

@hmgaudecker
hmgaudecker requested a review from timmens August 25, 2026 07:46

@timmens timmens left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: Carry the NNBEGM replay policy through a separate solve and simulate

Placement: src/lcm/model.py, Model.solve and Model.simulate

The new candidate replay works when simulate() performs a fresh solve internally.
The documented separate workflow still cannot use it:

values = model.solve(...)
model.simulate(..., period_to_regime_to_V_arr=values)

solve(return_simulation_policy=True) can return the required policy mapping, while
simulate() has no argument that accepts it. Supplying value arrays therefore leaves
period_to_regime_to_sim_policy=None and silently restores the foreign action-grid
maximization. The original witness remains unchanged:

                         narrow action grid   wide action grid
Solved value                 -7.8459              -7.8459
Simulated value              -9.9077              -7.8841
Chosen outer action            0.01                 5.00

Could simulate() accept the policy mapping already returned by
solve(return_simulation_policy=True) and pass it through the existing internal
channel? When caller-supplied values require an NNBEGM replay policy and none is
provided, simulation should raise clearly before entering the forward loop. The
regression should pass the exact (values, policies) returned by one solve and compare
that simulation with the automatic solve-and-simulate path. The current regression
asserts that policies exist, discards them, and triggers a second automatic solve.

A SolutionResult bundling values and policies would give the cleaner long-term API
and persistence story. The explicit policy argument plus fail-closed validation is a
bounded fix for this PR and preserves value-only solution reads for comparisons and
warm starts.

Blocking: Rank the reconstructed NNBEGM candidates by canonical Q

Placement: src/_lcm/simulation/simulate.py, _replay_nnbegm_candidates

Replay currently interpolates every candidate's stored actions and stored solve value,
selects the largest interpolated stored value, and evaluates canonical Q only for that
winner. Interpolation can change the relationship between the stored value and the
value attained by the emitted action. The selected action can therefore be strictly
dominated by another reconstructed candidate already present in the replay bank.

In the public smooth toy at (wealth, illiquid) = (1.0467, 2.04):

current replay:       target 2.8571, consumption 0.9067, value -18.8542
canonical bank winner: target 4.2857, consumption 0.6695, value -18.1265
independent oracle:    target 4.2857, consumption 0.7768, value -18.0981

Across 651 off-grid states, the interpolated-value winner and canonical-Q winner
differed 62 times. The largest attainable-Q gain was 0.7277.

Could replay canonically score every represented joint candidate, apply canonical
feasibility and finiteness, and then select the largest Q with the existing solve-order
tie rule? If no candidate is valid, the current NaN action and -inf value sentinels
can remain. The simulation action grid should not become a fallback because it
represents a different candidate set.

The existing candidate batching machinery can bound this calculation. This also
matches the canonical associated-action/value principle used by the earlier nested
replay implementation. Once this is correct, replacing candidate_value with a
representation mask could be considered separately to reduce policy storage.

Blocking: Make the NNBEGM outer-action inversion contract explicit

Placement: src/_lcm/solution/nnbegm.py,
_NNBEGMPeriodKernel._candidate_outer_actions

NNBEGM solves in outer post-decision target space and later recovers the user's outer
action by evaluating the target function at actions zero and one and treating their
difference as a global affine slope. OuterContinuousMargin currently declares no
affinity or inverse requirement.

A valid nonlinear parameterization such as

def new_illiquid(illiquid, investment):
    return illiquid + investment**3

solves with finite values because the kernel binds new_illiquid directly. During
simulation, the affine reconstruction fails its round-trip, every affected adjuster is
converted to NaN, and the tested subjects silently keep their current illiquid stock.

Could we choose and enforce the supported contract here? Two coherent possibilities
seem available:

  1. Require NNBEGM's outer target map to be affine in the outer action with nonzero
    slope, conditional on its other inputs. Check that the function depends on the
    action and raise an actionable error whenever the existing forward round-trip
    fails.
  2. Support general invertible mappings through an explicit user-supplied inverse, for
    example an action_from_post_decision entry on OuterContinuousMargin, and verify
    its result through the forward map.

The earlier nested replay deliberately used grid fallback for a non-affine or
unresolvable mapping. That fallback is unsafe for NNBEGM because the regime action grid
and NNBEGM.outer_grid need not represent the same targets. Early rejection is a
sufficient bounded fix if general inversion belongs in a follow-up PR.

Regression cases should cover a cubic mapping, zero slope, and a supported affine map
such as new = old + 2 * action.

Discussion: Define the NNBEGM contract for Phased declarations

Placement: src/_lcm/solution/nbegm.py, model-stage declaration validation, and
src/_lcm/regime_building/processing.py, regime_declares_phased

There are two separate problems in the current phase handling:

  • Phased(solve=utility, simulate=utility) fails during model construction with an
    internal beartype violation because NBEGM validation passes unresolved Phased
    objects to a callable-only helper.
  • The replay eligibility scan omits koopmans_aggregator, so a genuinely
    phase-varying aggregator can receive an NNBEGM replay marker even though the replay
    contract requires phase-invariant aggregation.

The earlier EGM design deliberately retained grid argmax when phase variation changes
the simulation FOC. For NNBEGM, that grid path may search a different outer candidate
set from the solver. Could we therefore make the capability boundary explicit?

A bounded contract would be:

  • resolve declaration pools before NBEGM validation;
  • treat Phased(solve=f, simulate=f) as invariant when both sides are the same object;
  • include the Koopmans aggregator in phase-variation detection;
  • reject genuinely phase-varying NNBEGM models with a clear domain error until their
    simulate-phase inner and outer decisions can be replayed or reoptimized over the
    correct candidate set.

If genuine phase variation is intended to remain supported through grid maximization,
the PR needs to explain how the simulation grid is guaranteed to represent the
NNBEGM candidate set.

Non-blocking follow-up: Complete solution object

The replay work demonstrates that value arrays alone are not always sufficient to
reproduce a solver's forward policy. A separate PR could introduce a public
SolutionResult that bundles value functions and any solver-published replay artifacts
and is consumed directly by simulate() and persistence. The current value-only load
path should remain available for value comparisons and warm starts without loading
large policy banks.

The same interface question is already raised under
“Return and consume one complete solution artifact” in the review of PR 407.

@hmgaudecker

Copy link
Copy Markdown
Member Author

Thanks — the three blocking fixes outside the Phased item are now pushed on feat/nb-egm:

  • Separate solve/simulate: simulate(..., policies=...) consumes the mapping returned by solve(return_simulation_policy=True) and fails closed when supplied values need a missing replay policy (6eef61f9, with the fp32 inversion tolerance follow-up in cbecc3a2).
  • Canonical ranking: replay scores every represented, feasible reconstructed joint candidate by canonical Q before applying the existing first-in-bank tie rule (e699cd3a).
  • Outer inversion: NNBEGM now rejects nonlinear and zero-slope target maps and verifies supported affine recovery (9df8d622).

The Phased contract is being handled in its separate bounded repair. The complete-solution-object follow-up is tracked in #428. Focused fp32/fp64 batteries are green; the canonical-ranking GPU gate was also green (-4.07% warm runtime, 0% peak-device-memory change).

@hmgaudecker

hmgaudecker commented Aug 26, 2026 •

Copy link
Copy Markdown
Member Author

Final follow-up on the Phased item: this is now closed.

  • NNBEGM-family validation consumes one normalized solve-phase declaration pool.
  • Phase variation is classified by exact object identity across functions, states, state transitions, the regime transition, and the Koopmans aggregator.
  • Bare declarations and Phased(solve=f, simulate=f) are accepted when both fields contain the same object.
  • Genuine phase variation and carried-only states are rejected during Model(...) construction, before period-kernel construction or replay-policy publication.
  • The follow-up architecture repair carries the normalized solve functions and phase-variation paths in the solver contexts, so solver runtime no longer imports regime-declaration topology.

The bounded repair landed in d2c78143; the final context-boundary repair is fa85227f. Verification includes 49/49 focused tests at fp64 and fp32, 86/86 bounded integration tests, hooks and whole-tree typing, and the final cross-platform CPU, fp32/fp64 GPU, notebook, and typing CI—all green.

@timmens timmens left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Only one last comment!

Blocking: Reconstruct NNBEGM outer actions at the realized state

Placement: src/_lcm/solution/nnbegm.py, _candidate_outer_actions, and
src/_lcm/simulation/simulate.py, NNBEGM candidate replay

The new inversion guard proves recovery only at solve-state grid nodes. Simulation
later interpolates the recovered actions. Inversion and interpolation do not generally
commute, so an accepted conditional-affine map can emit a target outside the candidate
set that NNBEGM ranked:

def new_illiquid(illiquid, investment):
    return illiquid + (1 + 0.1 * illiquid) * investment

This map is affine in investment with a finite, nonzero slope conditional on its
other input, so it satisfies the contract stated by the new error. The public
two-period toy accepts it, and its solved value arrays are exactly equal to those under
the baseline parameterization. At the off-grid state
(wealth, illiquid) = (1.0467, 2.04), replay produces:

NNBEGM target selected during solution: 4.285714
exact action at the realized state:      1.865211
interpolated action used by simulation:  1.901299
target actually reached:                 4.329164

The resulting target lies outside both the keeper state and the declared
NNBEGM.outer_grid candidates, which violates the candidate-faithfulness invariant
asserted by the existing simulation regression. The reported Bellman value also changes
from -18.126415 to -18.104359, even though the two parameterizations produce
identical solved-grid values.

Could the replay payload retain each candidate's target identity, or enough metadata to
reconstruct it, and derive the outer action at the realized simulation state before
canonical-Q scoring? The replay path can then apply a forward round-trip check and
score the action that actually realizes the solve candidate. The keeper/outer-major
ordering and existing tie rule can remain unchanged.

If realized-state inversion is outside this PR's scope, a bounded alternative is to
define and enforce a narrower contract under which the inverse action surface is
represented exactly by replay interpolation. A constant action slope alone is
insufficient when the affine intercept is nonlinear in another interpolated state; a
jointly affine target map over the continuous replay coordinates with a
state-independent action coefficient would be one simple sufficient restriction.

A public solve-to-simulate regression could use the function above and assert in fp32
and fp64 that every emitted target remains either the realized keeper target or an
NNBEGM.outer_grid node.

@hmgaudecker

Copy link
Copy Markdown
Member Author

Fixed in abe354de. The replay payload now retains each candidate’s outer-target identity instead of its solve-grid action (so there is no additional full candidate bank), resolves the simulate-phase target DAG per period, reconstructs the outer action at the realized state, forward-round-trips it, and only then applies canonical-Q ranking. Keeper/outer-major/discrete order and first-tie behavior are unchanged.

The public state-dependent-slope witness now passes in fp32 and fp64; the final affected battery is 44/44 in each precision, with Ruff, formatting, full ty, and commit hooks green.

@hmgaudecker
hmgaudecker merged commit 8bfcada into main Aug 26, 2026
23 checks passed
@hmgaudecker
hmgaudecker deleted the feat/nb-egm branch August 26, 2026 14:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants