Collective / gated regimes + edge-fold of IID shocks - #409
Conversation
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
Documentation build overview
46 files changed ·
|
c2d8650 to
563d80b
Compare
Head-vs-base failure diff — this branch causes 8 failures, all in
|
| tree | result |
|---|---|
feat/continuous-outer @ db00f64 |
30 passed |
collective-regimes @ 5c0b829 |
8 failed, 22 passed |
There are two distinct causes.
(a) map_coordinates now returns a WRONG GRADIENT at every on-node coordinate — 1 test.
zero_safe_weighted_term(w, v) is w * where(w == 0, 0, v). Masking the value with a
hard select is sound when w is a constant — a Pareto weight, a regime-transition
probability, a quadrature weight. An interpolation corner weight is not constant: it
is a differentiable function of the very coordinate being differentiated. At an
exactly-integer coordinate one corner weight is exactly 0, where zeroes that corner's
value, and the term becomes w * 0 — whose derivative is w'(c) * 0 = 0 instead of
w'(c) * v. The gradient loses precisely the corner whose weight is changing.
Minimal witness, grid[i] = i² so the correct slope on [i, i+1] is grid[i+1] - grid[i]:
c=1.0 value=1.000 grad=-1.000 on_node=True (correct: 3.0)
c=1.5 value=2.500 grad= 3.000 on_node=False correct
c=2.0 value=4.000 grad=-4.000 on_node=True (correct: 5.0)
c=2.5 value=6.500 grad= 5.000 on_node=False correct
c=3.0 value=9.000 grad=-9.000 on_node=True (correct: 7.0)
d/dw [zero_safe(w, 7.0)] at w=0.00 -> 0.0 (should be 7.0) <-- the hard select
d/dw [zero_safe(w, 7.0)] at w=0.25 -> 7.0
d/dw [zero_safe(w, 7.0)] at w=1.00 -> 7.0
Off-node gradients are exact; on-node ones come back as -grid[c], i.e. only the
surviving corner's -1 * v_lo term. Values are unaffected — this is gradient-only, which
is exactly what makes it easy to miss. test_gradients[map_coordinates] catches it:
-10.0 where 2.0 is expected.
(b) The helper rejects integer weights — 7 tests.
BeartypeCallHintParamViolation: parameter weight="JitTracer(int32[7])"
violates type hint FloatND | ScalarFloat | float
The previous expression was a bare multiply with nothing to violate. All seven are the
int32 parametrization of test_map_coordinates_against_scipy /
test_map_coordinates_round_half_against_scipy; the float variants pass.
Suggested direction
(b) is an annotation/dtype-coverage fix. (a) is the substantive one and I do not think it
should be fixed by loosening the test: the zero-safe contract as written is a statement
about values, and map_coordinates needs it to hold for derivatives too. Either
map_coordinates keeps the plain multiply (its corner weights are never ±inf-valued,
so it may not need zero-safety at all), or zero_safe_weighted_term needs a
custom-JVP/stop_gradient-free formulation that preserves ∂/∂w at w == 0. Worth
checking whether any other caller passes a weight that is differentiable in the argument
of interest.
Inherited from the base — 68
Unchanged from db00f64; none of these are this PR's doing.
| file | n |
|---|---|
tests/test_solution_on_toy_model_stochastic.py |
24 |
tests/test_solution_on_toy_model_deterministic.py |
8 |
tests/simulation/test_subject_batching.py |
7 |
tests/solution/test_n_nbegm.py |
6 |
tests/solution/test_n_nbegm_finite_baseline.py |
6 |
tests/test_processes.py |
5 |
tests/solution/test_negm_bequest.py |
3 |
tests/solution/test_negm_serviceflow.py |
3 |
tests/solution/test_egm_carry.py |
1 |
tests/solution/test_egm_discrete.py |
1 |
tests/solution/test_n_nbegm_continuous.py |
1 |
tests/solution/test_n_nbegm_epstein_zin.py |
1 |
tests/solution/test_nbegm_distributed_carry_sharding.py |
1 |
tests/test_regression_test.py |
1 |
The seven test_subject_batching.py failures are worth calling out: they sit in
tests/simulation, which this PR modifies heavily, so they were my prior candidate for a
genuine regression. They fail identically on the base.
Method, and what is NOT covered
One process per chunk (peak RSS resets between them); -q -n 2 --tb=no -rf, with
### EXIT <chunk> = $? recorded after each so a chunk killed by the memory ceiling shows
up as unmeasured rather than silently counted as clean.
The four
tests/test_mahler_yum_*.pyfiles are UNMEASURED on both sides — that
chunk exited143(SIGTERM from the memory cap) on head and base, even at-n 1.
Symmetric, so it is unlikely to hide an asymmetry, but I have not measured it and am
not claiming it is clean.
The base was measured in a fresh detached worktree at db00f64, not in the shared
one: that worktree had uncommitted work in src/_lcm/egm/upper_envelope/query.py, the
same area as most of the inherited failures, so measuring there would have compared this
PR against someone's work in progress. _lcm.__file__ was checked to resolve inside each
tree, since a linked pixi env can otherwise resolve the package to a different branch.
Fixed —
|
head 9f75cc6 |
base db00f64 |
|
|---|---|---|
| failures | 68 | 68 |
| caused by this PR | 0 | — |
Same caveat as before: the four tests/test_mahler_yum_*.py files remain unmeasured on
both sides (that chunk hits the memory ceiling even at -n 1), so this is "no
difference in what was measured", not "clean everywhere".
|
@/tmp/claude-1000/-home-hmg-econ-dev-pylcm-lcm-replications-src-lcm-reps-ecksteinCareerFamilyDecisions2019/da153776-cb2a-4e8a-9fc3-5b37c71b1f06/scratchpad/pr409-r3.md |
`trim_pad_from_raw_results` enumerated the fields it knew about -- `V_arr`, `actions`, `states`, `in_regime` -- which is a second, silent copy of `PeriodRegimeSimulationData`'s definition. It went stale: `nested_policy_fallback` was added to the structure without being added to the list, so `dataclasses.replace` carried the untrimmed field straight through. With 7 subjects at batch size 2 the leaf came out 24 rows against a 21-row `_in_regime` mask and `to_dataframe` died in `column[mask]` with an IndexError. Attribution, measured in fresh detached worktrees rather than assumed: origin/main eeb35f5 10 passed feat/continuous-outer d83e75f 7 failed, 3 passed collective-regimes d82a490 7 failed (pre-merge; inherited) So the regression lives on feat/continuous-outer and #409 merely inherits it, and tests/simulation/test_subject_batching.py is byte-identical to main's -- an unchanged test catching a production change, not a test that drifted. The fix reads the field list off the dataclass, so the next field added cannot reintroduce the class. The regression test is written the same way: it builds and measures every field generically, guards the mapping-vs-array assumption it makes so it cannot go stale silently, and asserts it inspected something -- a coverage check that quietly examines nothing would otherwise pass forever. Verified discriminating by re-introducing the exact enumeration: the new test and all 7 original failures come back, and both go green again on the fix. tests/simulation + tests/regime_building: 476 passed, 2 skipped (was 463 passed, 7 failed). ruff, ruff-format and ty all clean. This belongs upstream on feat/continuous-outer; it is landed here first because #409 is red on it. The contouter worktree is in active use by another session, so the cascade is theirs to take -- this merge becomes a no-op once it does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M4Vd9xVvBDhAzoxqGoriPn
`trim_pad_from_raw_results` enumerated the fields it knew about -- `V_arr`, `actions`, `states`, `in_regime` -- which is a second, silent copy of `PeriodRegimeSimulationData`'s definition. It went stale: `nested_policy_fallback` was added to the structure without being added to the list, so `dataclasses.replace` carried the untrimmed field straight through. With 7 subjects at batch size 2 the leaf came out 24 rows against a 21-row `_in_regime` mask and `to_dataframe` died in `column[mask]` with an IndexError. Attribution, measured in fresh detached worktrees rather than assumed: origin/main eeb35f5 10 passed feat/continuous-outer d83e75f 7 failed, 3 passed collective-regimes d82a490 7 failed (pre-merge; inherited) So the regression lives on feat/continuous-outer and #409 merely inherits it, and tests/simulation/test_subject_batching.py is byte-identical to main's -- an unchanged test catching a production change, not a test that drifted. The fix reads the field list off the dataclass, so the next field added cannot reintroduce the class. The regression test is written the same way: it builds and measures every field generically, guards the mapping-vs-array assumption it makes so it cannot go stale silently, and asserts it inspected something -- a coverage check that quietly examines nothing would otherwise pass forever. Verified discriminating by re-introducing the exact enumeration: the new test and all 7 original failures come back, and both go green again on the fix. tests/simulation + tests/regime_building: 476 passed, 2 skipped (was 463 passed, 7 failed). ruff, ruff-format and ty all clean. This belongs upstream on feat/continuous-outer; it is landed here first because #409 is red on it. The contouter worktree is in active use by another session, so the cascade is theirs to take -- this merge becomes a no-op once it does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M4Vd9xVvBDhAzoxqGoriPn (cherry picked from commit 7916291)
Absorbs upstream #407, which itself carries the #406 perceived-stochastic- transitions merge (c0699e4). The substantive upstream change for this branch is the age-normalization refactor: `_resolve_age_specialized_state_grids` became `normalize_age_specialization` + `grid_schedule`, `resolve_specialized_nodes(x, age)` became `resolve_periodized_nodes(x, representative_period)`, and `_grid_identity` / `_node_fingerprint` / `_continuation_grid_signature` are gone. Conflicts resolved by provenance, not by side-taking: * `regime_building/processing.py` (8 conflicts) — took upstream's normalization pipeline wholesale; kept `regime_to_v_interpolation_info_for_Q`, the collective E1 `compute_intermediates` guard, the `get_Q_and_F_collective` branch and the four gated-edge fold helpers. The collective branch is now ported to `resolve_periodized_nodes(..., representative_period)` so it resolves functions and constraints exactly as the singleton branch does. * `regime_building/Q_and_F.py` — upstream widened `QAndFFunction.next_regime_to_V_arr` from `FloatND` to `MappingProxyType[RegimeName, FloatND]`. Auto-merge updated the singleton builders only; the two collective twins (lines 735, 1285) were left behind and are fixed here. `ty` caught this, not the test suite. * `tests/regime_building/test_collective_phase_role_class_invariant.py` — the test pinned the literal `resolve_specialized_nodes(continuation_functions, age)`, which upstream renamed. Production was correct; the assertion is re-expressed through a derived `_resolver_name()` helper so it tracks the resolver rather than its spelling. Attribution against the c0699e4 baseline (fresh worktree, PYTHONPATH override, `_lcm.__file__` asserted, same pixi env): * `tests/solution` — 1212 passed, 3 skipped, 1 xfailed, 0 failed. * rest of `tests` — 38 failed + 7 errors, an EXACTLY identical set at the baseline (`comm` empty in both directions). All 38 are one upstream defect: d0d655c added the `nested_policy_fallback` simulation column without regenerating the pickled fixtures, so every failure is a one-column DataFrame shape mismatch. The 7 errors are the separate AdaptiveOuterMesh OOM. No new failures are introduced by this merge. Note for anyone attributing failures on this stack: `beartype_package()` caches INSTRUMENTED bytecode as `__pycache__/*.opt-beartype<ver>.pyc`, keyed on the beartype version and not on the conf. A merge can leave a `.pyc` compiled under a different conf — `is_pep484_tower` off — which is NEWER than its source, so no mtime check catches it. That produced 120 phantom `tests/solution` failures here (`int 0 not instance of float` in code byte-identical to the passing baseline). Clear `src/**/__pycache__` before believing any such result. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01M4Vd9xVvBDhAzoxqGoriPn
|
Cascaded
The 38 DataFrame-shape failures this branch was carrying are gone. They came from The 7 remaining errors are #407's, not this branch's
This would hit CI too. Both are This looks like the class an in-flight external audit already has open (workspace proportional to Note for anyone verifying locallyDo not try to phase |
Brings the 390->400->405->406->407 cascade to #409: ffb25e2, ec96c12, af4b05b, 06bf038, a60a092 and cdd4f65 are now all ancestors. Both conflicted files carried real 407-side fixes, so neither side could simply win and `-X ours` would have discarded them silently. Stateless targets (ec96c12: a target carrying no state is worth its utility, not zero) now seed `mixture_terms` instead of being accumulated as a raw `E += p*V`. The incoming raw product would have reintroduced `0 * -inf = NaN` here, where dissolution makes `-inf` continuations ordinary; routing them through `_sum_regime_mixture` keeps the zero-mass masking and the value-ordered, alpha-renaming-invariant reduction. `_sum_regime_mixture` accordingly lifts rank-deficient values to the common value shape before stacking, since a stateless target's V is rank-zero while a carry target's contribution carries the cell and stakeholder axes. The `len < len` guard covers the all-stateless case, where the values would otherwise right-align against the probabilities and weight a cell axis by the target axis. `reachable_targets` unions 407's declared-target keys with the existing per-target and coarse-candidate sets rather than being replaced by them. Replacing would have bypassed `_fail_if_coarse_candidate_folds_ambiguously` and the fold-round4 F2 guarantee. The union is safe in the direction that matters: a spurious candidate carries zero transition probability and contributes nothing, whereas omitting a routed one silently drops a continuation. Restores the `QNAME_DELIMITER` import that 407 dropped along with its only use. Verified: ty clean; tests/regime_building 368 passed; tests/simulation + stateless-continuation 142 passed, 2 skipped, 6 xfailed. The one failure, test_simulate_constraint_reads_next_state.py::test_the_constraint_is_available_as_an_additional_target, is pre-existing and stack-wide -- it fails identically on 18f5491 (#406), 0e3a574 (#407 pre-cascade), f7aa590 (#407) and 28d05bb (#409 pre-merge), each checked in an isolated worktree with _lcm.__file__ asserted inside the tree.
Brings the combined cascade hop to #409: bf61b14 (#400), 8f31991 (#406 transition-dependence fix), ddbec5b and 5b4339e are now all ancestors. One conflict, in `Q_and_F.py`, and both sides were right about different things. #407 `d9aec7d` transforms a stateless target's value BEFORE the expectation: `QuasiArithmeticMean` is `CE = g^-1(sum_r p_r * E_w[g(V'_r)])`, so `transform` precedes every expectation including the regime-transition one, and `inverse` applies once to the finished sum. Previously a stateless target was probability-weighted raw and only then inverted, leaving the transformed domain. #409 routes stateless targets through `mixture_terms` rather than a separate `E += p*V`, so they get `zero_safe_weighted_term` (a zero-probability +-inf is masked to exactly 0 instead of yielding NaN) and the value-ordered, alpha-renaming-invariant reduction. On this branch dissolution makes `-inf` continuations ordinary, so a raw product is a live NaN source, not a hypothetical. Resolved by doing both: CE-transform the stateless value, then seed `mixture_terms` with the transformed value. `_sum_regime_mixture` forms `sum_r p_r * g(V_r)` zero-safely and `ce.inverse` is applied once to the result, which is exactly `g^-1(sum_r p_r * g(V'_r))`. Also drops `test_collective_solve_default_call_still_returns_bare_mapping` (external test-suite dedup audit, round 1): it is a local mirror of `test_collective_regime_simulate.py::test_public_model_solve_default_return_shape_is_byte_identical`, which is the module it already imports its model helper from, and whose assertions are a strict superset. The file's own singleton guard, `test_singleton_solve_bare_call_returns_legacy_mapping_not_a_tuple`, is retained -- that is the invariant this module exists for. The other five deletions from that audit are shared with `main` and went to #390 as a patch. Verified: ty clean; tests/regime_building + the two touched test files 379 passed; the certainty-equivalent bearing files (test_temporal_aggregation, test_dcegm_validation, test_nbegm_epstein_zin_composite_flow, test_simulate_terminal_rows) 43 passed. KNOWN COVERAGE GAP, not closed here: no test appears to construct a model with BOTH a certainty equivalent AND a stateless target, so the intersection this resolution creates is exercised by neither side's tests. `d9aec7d` itself shipped without tests. The two halves are covered separately; their combination is not.
`d9aec7d` ("Transform stateless targets before the certainty-equivalent
expectation") shipped without tests, and no test in the suite built a model
carrying BOTH a certainty equivalent and a stateless target -- each half was
covered, the intersection was not. Patch authored by the #409 session while
merging 407->409; it belongs here so the fix and its test cascade as a unit.
`QuasiArithmeticMean` is `CE = g^-1(sum_r p_r * E_w[g(V'_r)])`, so `transform`
applies before EVERY expectation, the regime-transition one included. A target
carrying no state has no stochastic expectation to take but is still inside that
sum, so it must be transformed too. `_next_regime` routes to `dead` with
probability 1, which makes the identity sharp and independent of the mixture
arithmetic: with a single target at `p = 1`, `g` and `g^-1` cancel exactly, the
continuation is `V_dead` for ANY `g`, and the solved value is the closed form
`max_{c <= wealth} log(c) + discount_factor * V_dead`.
It discriminates, verified here rather than taken on the author's word. At
`d9aec7d`'s parent `3934701`, tree-isolated with `_lcm.__file__` asserted inside
the extraction, it fails at all three risk aversions; at `risk_aversion = 2` the
solved values are `[-0.218147, 1.486601, 2.084438, 2.084438, 2.084438]` against
the correct `[1.206853, 2.911601, 3.509438, 3.509438, 3.509438]`, reproducing the
reported numbers exactly. The gap is a constant `1.425 = 0.95 * (2.0 - 0.5)` on
every element, i.e. the continuation arrived as `0.5` rather than `V_dead = 2.0`,
and `g^-1(2) = 2^-1 = 0.5` at that risk aversion -- so the failure is
specifically the "left the transformed domain" defect, not merely a wrong value.
Tests only, no `src/` change. Full file 34 passed (was 31); hooks clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
Completes the cascade #405 -> #406 -> #407 -> #409, carrying `feat/nb-egm` @ `48d3795c` and with it the aggregator specification objects, the EGM preference-map refactor, and the runtime inactive-target check. Fourteen conflicts, and unlike the branches above this one they are not all import lists. `W_linear` is gone, replaced by the `LinearAggregator` dataclass. This branch had 88 references across 14 test modules, all of the form `koopmans_aggregator=W_linear`, plus one docstring. Migrated to `LinearAggregator()`. These would have failed loudly at import, so they were never the dangerous half; the dangerous half is the identity check that silently stops gating, and there is none left -- no `W_linear` or `W_epstein_zin` survives in `src/` or `tests/`. Certainty-equivalent reduction (`Q_and_F.py`), the substantive one. This branch collects UNMULTIPLIED `(regime, p, E[V])` terms and contracts them ONCE through `_sum_regime_mixture` -- value-ordered, and deliberately not keyed on regime labels (round-8 accuracy, round-10 F1). Upstream folds `CE = CE + p*E[V]` per target and divides by the represented mass. Kept both: the single contraction produces `CE`, then upstream's mass gating applies to it unchanged (`_unit_regime_mass_or_nan` on the per-target route, `_regime_mass_is_unit` selecting against NaN on the lottery route). These compose exactly because `_sum_regime_mixture([])` returns `zeros_like(like)`, which is the zero upstream's accumulator started from. `_scalar_target_contribution` now returns the mixture terms AND the probability mass, rather than one or the other. The reducer stays `zero_safe_average`, NOT upstream's new `_expectation_over_stochastic_nodes`. The latter guards the weight SUM only, so `0 * -inf` still poisons its numerator -- and `-inf` continuations are routine here, where dissolution makes them ordinary rather than exotic. The PREDICATE does move to upstream's `continuation.has_lottery_axes`, since the `next_V_has_stochastic_states` mapping it replaced no longer exists and would have been a `NameError`. `_get_joint_weights_function` changed contract upstream: it takes the ordered tuple of lottery variables instead of re-deriving them, so that the weight axes and the productmapped value-surface axes are fixed by one ordering. Upstream updated the singleton path only; the collective builder is this branch's and still called the old signature. `ty` caught it. Updated to derive `lottery_variables` exactly as `_build_target_continuation` does. Duplicate-parameter trap, three times: `source_regime_name` is passed at the top of both `_process_regime_core` call sites and appears again in upstream's added block, and is declared twice in the signature -- which is a `SyntaxError`, not a type error. Kept one of each and took only `phase_name`. `ISC004` fires on six of this branch's error strings under the ruff config that arrived with the merge; wrapped each in parentheses rather than silencing it, since every one sits in a returned collection where a dropped comma would fuse two distinct messages. `PLR0915` then fired on `_process_regime_core` (55 > 50) because the merge added statements; extracted `_partition_phase_functions`, which is a pure move -- `all_functions` and the stochastic-name set were already local to that block. One real defect surfaced, the known coarse-law witness from `BUG-gridsearch-drops-inactive-target-mass.md`: `test_coarse_self_transition_retains_the_self_continuation` routed its entire mass to `stay` at `stay`'s last active age, where `stay` is inactive next period. It solved to an all-NaN value function and said nothing, because the test ran at `log_level="off"` where the validation is skipped. The coarse law is now age-dependent and routes to `done` at that age; period 0 -> 1 still exercises the F1 regression this test exists for, namely that a coarse law returning its OWN regime keeps the self-continuation. Converted to `log_level="debug"` per the fleet NOTICE, so the check actually runs. 748 passed across `tests/regime_building/` plus the aggregator tests, none on the AVOID list, `_lcm` asserted to this worktree on the controller and both xdist workers. `prek run --all-files` clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #409 +/- ##
==========================================
+ Coverage 92.05% 92.60% +0.55%
==========================================
Files 310 331 +21
Lines 32809 37503 +4694
==========================================
+ Hits 30201 34729 +4528
- Misses 2608 2774 +166
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Leg 4 (final) of the dcegm review-repair cascade. Seven conflicted files. Four were unions of two independent additions; three needed a decision. Mechanical unions (both sides kept): - `.pre-commit-config.yaml` — `name-tests-test` exclude is now the three-way union: `tests/mock_regime.py|tests/collective_fixtures.py| tests/envelope_configs.py|^tests/test_models/|^tests/data/|^tests/solution/_`. Both fixture modules exist in the merged tree. - `src/_lcm/model_processing.py`, `src/lcm/__init__.py` — import and `__all__` collisions from alphabetical interleaving. `ExactEnvelope` was already imported at line 165, so the `__all__` entry it gained is backed. - `src/lcm/model.py` — `_solve_compiled` now takes BOTH `collect_simulation_policies` (from #407) and `retain_dissolution_flags` (from #409); its body already called `solve()` with both. The auto-solve comment block keeps #409's dissolution-flag note and #407's two-signal note, since they document different things. Decisions: 1. `src/_lcm/simulation/simulate.py` — the two branches widened the SAME function differently. #409 returned `(actions, nested_fallback)`; #407 returned `(actions, reported_value, nested_fallback)`. The auto-merged return paths had already taken #407's 3-tuple, so the conflicted paths were resolved to match it, and the docstring to #407's three-item form (#409's fallback description is a subset of it). Verified afterwards: the signature and all six return paths are 3-tuples. 2. `src/_lcm/solution/backward_induction.py` — #407 recomputed `base_state_action_spaces` inside `solve()` at the old position while #409 had MOVED that computation earlier in the same function, so that a `_reject_edge_fold_state_param_collisions` fence could run before any kernel is compiled. Textually these are additions at different offsets, so the merge offered both and neither side conflicts — the relocated-refactor duplicate. #407's copy is dropped; #409's earlier one is kept, which preserves the fence's position. Kept #407's `host_device` gating (`None` unless policies are collected) and merged the two comments. 3. `src/_lcm/regime_building/processing.py` (7 hunks) — #407 renamed `is_egm_carry_target` → `continuation_demanded` and replaced the inline `egm_child_regimes` comprehension with the `_continuation_targets` helper. Took the rename at all six sites and the helper call, while keeping #409's four blocks in the same region (`fold_only_regimes`, the same-period-ref validation pair, the gated-edge validation, the folded-endpoint check), which #407 has no counterpart for. `is_egm_carry_target` and `egm_child_regimes` now appear nowhere in `src/` or `tests/`. `solver_kernels.continuation_template` → `.continuation_spec` came with the same wave; `continuation_template` survives as the delegating property on `engine.py:377`, so the two forms are equivalent. Verification: `ty` green; all six resolved modules parse; no conflict markers remain anywhere in the tree. Behavioural smoke, all on CPU: - `tests/simulation/test_egm_sim_policy_value_read.py` + `test_egm_associated_policy_read.py` — 5 passed in 1.20s - `tests/simulation/test_simulate_n_nbegm_continuous.py` — 5 passed in 146.02s (the file whose four tests the #407 leg broke and repaired) - `tests/solution/test_collective_gated_edge_age_specialized_axes.py` + `tests/simulation/test_aot_collective_and_gated.py` — 7 passed in 15.02s (this branch's own feature, the part decisions 2 and 3 could have broken) Known gap, cascaded in from upstream and NOT introduced here: this wave predates nb-egm `4db23f8b`, so all four legs carry the Windows-only exact-kernel break — `ModelInitializationError` raised where the conftest skip gate keys on `ExactAffineKernelUnavailableError`, 32 tests red on Windows only. The fix cascades as the next wave from nb-egm `dfa68097`.
…nto feat/perceived-stochastic-transitions
…tions' into feat/continuous-outer
The adaptive reader refuses a policy proposal whose recovered action lands beside a declared endpoint rather than on it. Control then reaches the grid baseline, which projects the raw pair onto each published mesh endpoint and admitted the result on domain containment alone. An action one representable step inside an endpoint satisfies containment, so the projection could publish the very stock the proposal had just refused, and a signed outer domain reaches that case for every subject whose stock and endpoint straddle zero. Each projection now reports whether its action reproduces the endpoint it aimed at, decided by the same tolerance-free predicate the proposal uses, and the baseline may rank a projection only where it does. The predicate reads the published mesh rather than the outer state's declared bounds, so a mesh that stops short of those bounds still has endpoints the check can recognise. The retreat to an endpoint's interior neighbour goes with it. It served subjects whose exact projection left the mesh, but it did so by aiming at a different stock from the endpoint being certified, which is the distinction the verdict now draws. The two endpoints are mirror images, so a subject missing the near one keeps the far one, which a sign-crossing action reaches exactly; a subject missing both leaves the read with nothing to publish and raises. The two-asset toy gains a terminal-utility override, since its default bequest is singular one unit into durable debt and a signed outer domain needs one that is finite there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The zero search-error bound both required numerical profiles carry is a claim about the solver's code, and nothing re-checked it after that code changed. It is now stated twice, in instruments that fail for different reasons. The structural half reads the enumeration route off the syntax tree of `grid_search.py` and `max_Q_over_a.py`: the whole action-name tuple reaches `get_max_Q_over_a`, the whole action mapping reaches the compiled core, the core product-maps over that tuple unbatched, and each of the four reductions -- the singleton and collective ones on both the solve and simulate sides, plus the taste-shock smoothing -- covers every action axis. Names and attribute chains are resolved rather than matched by line, so an edit that keeps the property keeps the test green. The executable half solves a small model once per shape and re-solves it from params, making each declared candidate in turn the only feasible one and, separately, the unique maximizer. A candidate the search never visits publishes -inf in the first sweep and a runner-up value in the second. The collective reduction is swept separately because it scalarizes on its own code. Two tests mask one cell inside the action product map and show the sweep going red on exactly that cell and green on the others, so a green sweep is evidence rather than a report that nothing raised. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gzx6oLKVkrBrxLXhseoznK
The candidate-set certificate proved exhaustiveness by solving. Simulation maximizes with `get_argmax_and_max_Q_over_a`, a different callable from the one backward induction reduces with, so a candidate the solve reaches is not thereby a candidate simulation reaches -- and a cell suppressed on the simulate side alone left every structural obligation green and every executable sweep green, because none of them simulated. Each of the three sweeps now runs again through `model.simulate`: the singleton unique-feasible sweep, and the two-stakeholder one, whose household argmax is a third reduction distinct from both the singleton simulate path and the collective solve path. Two further checks mask one cell inside the simulate reducer alone and require the published action to move on exactly that candidate and stay put on another -- the solve reduction is a different callable and keeps the full set, which is the divergence a solve-only certificate cannot see. One more obligation derives the certified-source set from the obligations themselves, so an obligation cannot come to rest on a source the contract does not name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gzx6oLKVkrBrxLXhseoznK
A rejected policy replay handed the fallback a choice between the two declared endpoints and nothing else. The published policy carries far more than that -- a keeper branch plus one conditional inner policy per outer mesh node -- so a subject whose recovered action could not reproduce the endpoint the solve ranked first was sent to the opposite endpoint even when a mesh node between them was both reachable and better. On the shipped three-node mesh `[-1, 4.5, 10]` that emits the far endpoint at canonical Q -9 while the middle node, which the subject reaches exactly, is worth -3.5. The fallback now ranks every branch the solve published. Each is admitted by the same tolerance-free predicate the on-mesh proposal uses, against the same declared inverse domain, so an unreachable branch removes that branch and nothing else; each surviving branch keeps its own inner action rather than another branch's; and the ranking is by the values the solve published, which is the comparison the solve itself made across these branches. The keeper occupies the leading row so it wins ties, matching the on-mesh convention that an exact keeper beats an equal-valued adjustment. The caller scores only the winner through canonical Q, which is what makes it comparable with the raw simulation-grid pair. Ranking a read whose leading axis is not the outer-node axis would pair every node with another node's policy, so the length disagreement is refused rather than ranked. The endpoint-only selector is deleted; nothing projects onto an endpoint any more. A subject reaches the refusal only when no branch at all is reachable, which needs a mesh whose every node is a declared endpoint and a keeper target outside the published domain. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The read hands its caller the branch a rejected replay falls back to, so the return description names it and says what the caller may do with it. The caller's own summary of the value it reports drops the word "projected": a rejected replay is given a published branch, not a pair projected onto an endpoint. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
The zero represented-candidate bound rested on a source set written down in four places with no join between them, and on an executable sweep that only ever made one candidate feasible at a time. Either could be defeated without turning anything red: drop a source from the contract and the digest checker still validated every digest it found while the certificate compared its own `_parse` literals to its own constant; suppress a candidate only when several are feasible and every one-hot sweep still passed while the simulated winner moved. `tests/candidate_certificate/sources.json` is now generated from the literal `_parse` call sites by AST, so the declared set is a consequence of the obligations rather than a constant beside them, and the certificate loads it instead of naming the sources itself. `verify.py` joins that inventory against the AST, the generated policy, an explicitly supplied profile contract, and the bytes on disk, demanding exact set equality in every direction; its `--self-test` seeds deletion, addition, duplication, rename and byte-change and requires each to be named. The two controls under `controls/` are repository-rooted ports of the mutations that found the gap, and each is red before the change and green after. Four new tests replace the one-hot family with the whole neighborhood: all 63 nonempty feasibility masks over the action product, for singleton and collective, solve and simulate, asserting the selected action and the published value against a scalar argmax that refuses ties and empty support and shares no control flow with the production reduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gzx6oLKVkrBrxLXhseoznK
Two changes that the pre-commit config forces into one commit: a hook rejects any commit made while `.pre-commit-config.yaml` is modified but unstaged, so the workflow bumps cannot land separately from the hook that gates them. Hooks. Five the config was missing: - `pixi-lock-check` at pre-push, so a stale lock fails here instead of burning a CI cycle. `--check` re-solves without writing, so it cannot dirty the tree. - `pre-push-hooks-installed`, which turns a clone that never ran `prek install -t pre-push` from a silent no-op into a visible failure. - `notebook-cell-source-format`, `no-hardcoded-user-paths` and `codespell`. The first two currently find nothing, which is the point of installing them before they would. codespell reported 103 hits and every one is correct as written. `statics` (comparative statics) accounts for 91, `arithmetics` and `disjointness` for another six; those are jargon and go in the dictionary, which now matches the shared one. Four are project-specific and take an inline ignore at their site so the word stays caught everywhere else: the MSVC object-output flag, which has to be spelled exactly; a deliberately misspelled key in the test that asserts unknown keys are rejected; and a regex prefix that has to match both singular and plural. Two prose uses of "re-uses" are simply reworded. CI versions. `pixi-version` was pinned to v0.76.1 while the pixi that writes `pixi.lock` locally is 0.78.0, so CI was validating a lock produced by a newer binary with an older one. pixi is largely forward-tolerant, so that fails quietly rather than loudly, and only stops working once the lock uses something the pinned version cannot read. `actions/upload-artifact` was three majors behind. v5 and v6 move the runtime to Node 24 and require an Actions Runner of at least 2.327.1; the five ubuntu-latest call sites are unaffected, and the two on `ubuntu-latest-gpu` will tell us on the next run whether that runner is current. `prefix-dev/setup-pixi` goes to v0.10.2. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
`pre-push-hooks-installed` and `notebook-cell-source-format` run scripts from the `.ai-instructions` submodule. pre-commit.ci clones without submodules, so both fail there with "Executable /code/.ai-instructions/hooks/... not found" rather than running. No workflow in `.github/workflows/` checks out submodules either, so these two can only ever run in a developer clone. Skipping them states that plainly instead of leaving two hooks permanently red. The first asks whether this clone installed its pre-push stage, which a runner has no stake in. The second genuinely loses CI coverage: notebook cell source formatting is now checked where the submodule exists and nowhere else. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
Neither describes how this project works. The plotting module says to use plotly and not matplotlib. No `.py` source file imports plotly — it appears only in notebooks and docs — and this file already states the same rule in its own Plotting section, so the include was carrying a duplicate of a rule the source tree does not exercise. The pandas module is about cleaning survey data: build the frame column by column, name each raw column at its assignment, reshape and merge only at the end. pandas appears in 39 files here, but as DataFrame construction for library output rather than as a cleaning pipeline; four files touch a CSV reader and none misuse `inplace`. Its version premise did hold — the environment resolves pandas 3.0.5 — so this is a relevance judgement, not a correction. The jax module stays, and earns it: `custom_jvp`/`custom_vjp` occur 24 times in `src`, `jax_enable_x64` 57 times across `src` and `tests`, and `register_pytree_node` 7 times. Its one piece of unused advice is the argument for keeping it rather than against — `lax.cond` and `lax.switch` appear zero times, so the warning that `jnp.where` evaluates both branches and can manufacture a NaN is guidance this code has not yet taken up. Also advances the .ai-instructions pointer to pick up the codespell dictionary and the pre-commit.ci skips for the submodule-backed hooks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
#409's base had moved five commits ahead, leaving the pull request `CONFLICTING` / `DIRTY`. GitHub could not build a merge ref for it, so pre-commit.ci reported "error during mergeable check" and no pull-request workflow ran at all — the branch looked untested rather than failing. The only content conflict was the `_replay_continuous_outer_decision` docstring, where both sides had rewritten the return description. The signature returns four values, so the four-item list is the accurate one; the sentence about inference refusing on a fallback entry is kept from the other side rather than dropped. `feat/continuous-outer` also installs the hooks that live in the `.ai-instructions` submodule and adds codespell. Three sites here needed attention: the two deliberate `probabilit` regex prefixes take the inline `codespell:ignore` this repository already uses for that shape, and one docstring now spells "redeclares". That last one moves a certified source, so the candidate-certificate inventory and the profile contract were regenerated from it — which is the architecture working: a one-word docstring edit turned both `generate_sources.py --check` and the unified verifier red before either was told anything had changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gzx6oLKVkrBrxLXhseoznK
… bank When a refined nested proposal cannot be replayed at a subject's realized state, simulation falls back to the branches the solve published: the keeper plus one conditional inner policy per outer mesh node. It now emits the canonical-Q maximizer over the branches that are both reachable and canonically feasible there, instead of the branch carrying the highest value the solve happened to publish for it. Three things change together, because each alone leaves a way for a valid published branch to be discarded. The selector returns the complete bank rather than one winner. Its rows are node-major, keeper first, each carrying its own conditional inner action, and its mask covers only replay structure, interpolation support, finiteness and the solver-owned consumption budget. `_nested_grid_baseline` then scores every surviving row through canonical Q in one vmapped call and drops infeasible or non-finite rows before the argmax. The winner used to be chosen first and scored afterwards, so a branch the realized state makes infeasible could take the selection and then be rejected, discarding the rest of the bank with it. Ranking on published values could also prefer a branch that canonical Q ranks lower. Both the bank predicate and the per-pair admissibility check now read one domain, the published outer mesh. They used the declared inverse domain and the mesh extrema respectively, so a keeper landing in the gap between a wide declared domain and a narrower mesh was admitted by the first and rejected by the second, suppressing a reachable node that nothing reconsidered. Tie order is explicit: the keeper is row zero and wins an exact tie, then ascending mesh-node order; the grid baseline wins an exact tie against the best published row, preserving the no-degradation convention. When no branch survives the caller raises before emitting an action. The accompanying test module covers a lower-published-value feasible branch, a published/canonical order reversal, both tie conventions, inner-policy alignment, an all-infeasible bank, and a mesh narrower than the declared domain, each at both precisions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
…rdering A compiled profile policy carries one path and one digest per profile, so a certified set declared under `additional_sources` never reached the policy the protocol actually enforces. The contract's primary `source`/`source_digest` pair now names the generated inventory and its bytes, which every certified source feeds, so the compiled policy binds the whole set: change either source and the inventory moves, and the stale policy is rejected. `verify.py --policy` reads both the rich and the compiled projection and joins the anchor exactly. The executable half swept only the feasibility mask, at one fixed ascending ordering. Under that ordering the winner of a mask is its highest feasible index, so candidate 0 wins no multi-feasible mask and only the last candidate is ever the all-feasible winner — an omission of any other candidate changes nothing the certificate observes. `rank_vectors` generates one strict total order per candidate, each won by that candidate, and all four routes now enumerate every ordering against every nonempty mask. The AST gate requires both matrices and both loops, so removing the sweep is refused rather than silent. mt7 seeds a byte change in each certified source, an anchor replaced by each, and an absent anchor; mt8 collapses the ordering matrix and removes the sweep from each route, and exhibits the miss the fixed ordering admits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Gzx6oLKVkrBrxLXhseoznK
Selection now ranks the keeper and every published node by canonical Q, so the branch values the fixture published no longer decide which branch a fallback emits. The module steered selection with those values, which left ten tests asserting the superseded rule and three more passing whichever rule ran: the stub objective `-|investment|` makes the keeper the unconditional argmax. Each test now names the outer action its objective favours, and a refusal is made observable by peaking the objective at the branch the rule must refuse -- without that the favoured branch loses on value anyway and the test reports the same outcome whether the rule excluded it or not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Akf3HYBxuhGjKftH8jofDq
|
Cascaded Round 6 came back: eight of nine closedF1-F8 are Round 6 named two residual escapes. Both were reproduced here from the shipped
What changedThe contract's primary The certificate now sweeps the ordering as well as the mask: Two controls ship with it, each requiring all five of its perturbations to be rejected: Evidence at
|
| precision | tests | passed | failures | errors | skipped | time |
|---|---|---|---|---|---|---|
| 64 | 7090 | 7028 | 0 | 0 | 62 | 1193.6 s |
| 32 | 7055 | 6732 | 0 | 0 | 323 | 779.8 s |
Test and skip counts are identical to round 6 at both precisions, which is expected:
this round changed the content of four tests, not their number. The candidate
certificate is 50 passed / 0 failed at both precisions, and prek --all-files is clean
(34 hooks, 0 failed).
One thing that moved the wrong way
The protocol's own anchor checker now reports all_structural_markers_present: false
where round 6 reported true. This is a direct consequence of the anchor change: that
tool greps the file named by candidate_source for full-grid enumeration markers, and
that file is now a JSON inventory. Nothing about the production sources changed. The
enumeration route is instead resolved through the syntax trees of both certified
sources by the certificate's structural half, which is strictly more than a marker grep
over one file — but the tool prints false, and that is recorded rather than
paraphrased away.
The earlier OOM report in this thread is from 2026-07-29 and describes a tree many
commits behind; treat this comment as superseding it.
|
@timmens Final close-out: the current head is The last residual item was certificate architecture, not a demonstrated production error. The final form proves exact Round-9 verification at The follow-up head This supersedes my Round-7 status comment above. The follow-up to review 5064008752 was implemented at |
Windows checkouts use CRLF, so raw-byte digests rejected semantically identical certified sources. Hash canonical UTF-8 text while retaining sensitivity to non-newline changes.
timmens
left a comment
There was a problem hiding this comment.
Thanks for the substantial follow-up. I reviewed the current head 1bff3e16b from
494a3a5c7, reran the previous-finding battery, and checked the declaration redesign
and candidate certificate. The earlier joint-transition, callable-probability,
zero-argument-weight, and duplicate-declaration-surface concerns are resolved.
I am approving conditionally on the required CI becoming green. Windows currently
fails because the candidate certificate hashes raw checkout bytes, so CRLF line endings
change the source seals; the benchmark job is still running. The remaining review
points below are non-blocking, and I am happy for you to decide whether they belong in
this PR or a follow-up.
Non-blocking correctness follow-ups
Preserve ordinary DAG dependencies of carried-state Pareto weights
A Pareto weight may now read a carried state, which fixes the direct case. The carried
state's solve imputation is also a first-class regime function, so it may read another
function in the regime DAG. In a small public model I used weight_f(power),
impute_power(power_signal), and a normal regime function named power_signal. Model
construction succeeds, then solve(log_level="debug") raises
KeyError: 'power__power_signal'.
_compose_carried_imputations includes the carried imputations alone and consequently
reclassifies the ordinary function read as a parameter. Could the weight evaluator
compose the reachable solve-function DAG while preserving the existing
function/parameter distinction? A regression test can assert the resulting stakeholder
values (3, 0). This is a specialized, immediately visible failure with a simple user
workaround, so I do not consider it a merge blocker.
Preserve narrow reachability for Phased target mappings
The folded-conditioner guard appears to lose the declared target set when the regime
transition is Phased. I reproduced this with two disconnected paths:
safe_entry -> folding holds risk_type fixed, while
moving_entry -> sideways changes it. moving_entry.transition is
Phased(solve={"sideways": ...}, simulate={"sideways": ...}). Model construction
rejects the valid folding path and says moving_entry reaches folding.
_reachable_regime_targets recognizes a direct mapping and treats the outer Phased
object as a coarse transition over every regime. Could it resolve either phase first?
The phase grammar already guarantees matching forms and target keys. This requires a
rare combination of features and fails during model construction, so I also consider
it non-blocking.
Small documentation cleanup
The single declaration surface looks much clearer to me. One small remnant remains:
the public Model.solve and Model.simulate docstrings describe dissolution in terms
of GatedEdge, while that class is now internal and users declare a
ValueDependentTransition. Could those two references use the public name and
vocabulary?
Focused local checks at this head: 64 previous-finding tests passed, 405 declaration
and integration tests passed, and all 93 candidate-certificate tests passed.
What this PR adds
This PR lets pylcm represent collective life-cycle decisions in which several
people choose one household action but retain distinct utilities, continuation
values, and outside options. It also adds two complementary ways to handle
uncertainty without unnecessarily enlarging stored value functions: folding a
purely transitory IID shock and declaring a correlated
JointTransition.The existing singleton-regime workflow remains the default.
Economic model
Collective choice
A regime becomes collective by declaring a
CollectiveUtilityin the slot italready has for utility. For each feasible state-action cell, pylcm constructs
one action value per stakeholder,
Q_s(x, a). The household chooses one commonaction
and stores every stakeholder's value
Q_s(x, a*(x))at that same choice. Thisis the key distinction from solving separate optimization problems for the two
partners.
Everything is declared in a slot the regime already has, so there is exactly one
way to write a collective regime:
The
utilitieskeys are the stakeholders, in insertion order, and that orderfixes the trailing axis of
Vand of every published array. A stakeholder whosebody is
Nonedelegates to autility_<s>entry arriving from the model level.Value constraints may read each stakeholder's
Q_s, which is what supportslimited-commitment applications where household decisions must satisfy
individual rationality rather than only maximize a household objective. When no
action satisfies all participation constraints, the cell carries a dissolution
flag;
solve(..., return_dissolution_flags=True)publishes those flags andsimulate(period_to_regime_to_dissolution_flags=...)takes them back.Consent, separation, and fallback regimes
ValueDependentTransitiondeclares a transition whose continuation depends on aBoolean economic condition — mutual consent before entering a couple regime, or
dissolution when a participation constraint fails. It goes in the
transitionslot, keyed by target:
The key is always the gate-open target. A dissolution edge is keyed by the
continuing collective regime under
gate = ~D_target; keying it by thesingleton would send both partners there whenever the couple stays together.
routesis keyed by source stakeholder, and each route owns four destinations:the open regime (the dict key) and role (
target_stakeholder), and the closedregime and role (
fallback.regime,fallback.stakeholder). If the gate is openthe leg reads the appropriate stakeholder value in the target regime; if closed
it reads the declared same-period fallback, usually that stakeholder's single
regime. The ordinary regime-transition draw still determines which edge is
attempted. Solve and simulation use the same routing declarations, including
target-specific state projections and stakeholder roles.
Simulation follows a fixed-size cohort of rows. A separation does not create two
new rows;
own_stakeholdersays which stakeholder each simulated row followsafter stakeholder-specific routing.
One declaration surface
An earlier revision of this branch shipped two public spellings for the same
regime: the declarations above, and the lowered form the engine runs after
pylcm takes them apart —
stakeholders,pareto_objective,value_constraints,same_period_refs,gated_edgesasRegimearguments, plus a publicGatedEdgeclass.Publishing both would have made the equivalence and collision rules between two
spellings part of the long-term public API. The lowered vocabulary has been
removed from the public surface:
field(init=False)derived values. They survive asoutputs of lowering, which is the supported way to ask what a declaration
produced, but they cannot be passed to the constructor or to
replace;functions,constraintsandtransition, so thedeclaration objects stay where the author wrote them, and the engine reads
three derived views instead (
decomposed_functions,decomposed_constraints,decomposed_transition);GatedEdgemoved to_lcm.gated_edgeand leftlcm.__all__.This costs no user. An AST walk over every
*.pyundertests/— countingast.keywordnodes named after one of the five slots, 0 parse failures at eachref — finds 0 occurrences at the branch point and 7 at the head, none of
them to
Regime(...). The lowered vocabulary never existed before this branch,so removing it before release breaks no promise that was ever made. The seven
survivors are engine vocabulary — five
_MockRegime(which bypasses__init__),one
build_pareto_weights, onedataclasses.replaceon a canonical engineregime — and each is pinned by name in a census test that fails if an unlisted
one appears.
Closing the lowered form also required widening the declarations to cover three
configurations that had been reachable only in lowered form: a gated edge inside
a
Phasedtransition, aPhasedper-stakeholder utility, and per-stakeholderutility bodies arriving from the model level.
Uncertainty and value-function size
Folded IID shocks
An IID process can set
fold=Truewhen its realization is observed, used in thecurrent decision, and then discarded. pylcm still solves the choice problem at
every quadrature node and then averages the resulting values,
E_epsilon[max_a Q(x, epsilon, a)], but does not store the shock as a value-function axis.
This is deliberately restricted to cases where the reduction is exact: a fully
specified, non-persistent IID process under linear expectation, solved by grid
search, with no later transition or same-period value read depending on the
discarded realization. A folded shock whose conditioning state can move between
periods is refused, because the fold would otherwise price a distribution the
declared transition does not use. Unsupported combinations fail during model
construction.
Correlated next-state innovations with
JointTransitionSeparate
MarkovTransitiondeclarations imply separate draws. That is wrongwhen one innovation jointly determines several next-period states—for example,
a partner offer carrying correlated earnings and health characteristics.
JointTransitiondeclares one finite-support lottery on a particularsource-target edge and several output laws that share the same support row.
Backward induction forms the expectation over that common draw inside the
source action value before maximizing, while simulation samples one row and
uses it for every output. The latent realization is transition-local: it is not
a model state, initial condition, simulation column, or stored value-function
axis.
Phasedcan give solving and simulation different perceived and realizedjoint laws while preserving the same structural support.
An output may land on a stochastic process's own nodes as well as an ordinary
grid, which is how correlated innovations reach a grid pylcm discretized rather
than one discretized by hand. Probability vectors are never silently normalized:
a lottery that does not return one probability vector per source cell is an
error, raised in the params-bound preflight.
Transition architecture
The branch removes the former
TransitionLawsmetadata representation. Ordinarydeterministic and Markov state laws, stochastic-process laws, explicit process
entry, interpolation bases, and
JointTransitionkernels now lower to onetarget-specific
TargetTransitionPlan. Solve, simulation, validation,diagnostics, and solver-scope checks consume that same plan.
Regime-transition probabilities remain a separate concern: they select the
target regime, whereas a target transition plan determines the target states
conditional on that target. Both public grammars are canonicalized by reachable
target, but combining target selection and conditional state evolution into one
public object would blur rather than remove this distinction. This supersedes
the internal-representation question raised in #423 without adding a wrapper
around
Regime.transition.Correctness and supported scope
The implementation also adds the less visible machinery needed to make these
features compose safely:
handling, and target-specific parameter namespaces;
incomplete producers;
0 * -inf -> NaNinregime mixtures, stakeholder aggregation, and gated fallbacks;
judged even when a weight reads a carried state, whose solve value comes from
its imputation rather than from a grid axis;
simulation; and
is not implemented.
Collective regimes currently use
GridSearchand linear expectation and do notcombine with folded shocks or EV1 taste shocks. Transition-local
JointTransitionlotteries likewise useGridSearch; EGM-family solvers rejectthem explicitly. A gated edge on a coarse or terminal transition is unsupported
and refused: a gate is a route, and a transition naming no target has no route
to fall back to.
The branch includes end-to-end tests of the documented public workflow and
class-level regression coverage for collective choice, participation and
dissolution, gated routing, folded-shock timing, joint-transition correlation,
parameter ownership, float32/float64 edge cases, and solve/simulation/AOT
parity.
Review status
Nine bounded external mathematical/code-review rounds are complete. The definitive
Round-9 closure audit returned
CLEAN / CLOSED: F1-F9 areverified_fixed, bothrequired profiles pass, no finding remains open, and
next_actionisdone.The final candidate certificate is an exact route-local direct-data-flow proof over all
six GridSearch corridors (ordinary and taste-shock singleton/collective solve and
simulation). It seals the 43-source dependency frontier and proves that the exact
Q_arrand
F_arrproduced byQ_and_Freach the full reducer and publication path without acandidate-changing transformation. The synchronized closure suites reject all 66/66
semantic mutations and all 176/176 full re-anchored mutations, including the historical
MT9 and MT10 escapes.
The follow-up for review 5064008752 is complete at
90fe284f. It preserves ordinaryDAG ancestors when carried-state Pareto weights are folded into solve-time imputation,
retains narrow reachability for
Phasedper-target transitions, and updates the publicModeldocstrings to nameValueDependentTransitionrather than the lowered internaledge type. This repair head also incorporates #407 at
1c9f1d22and re-anchors the sourceinventory and direct-flow seals to the combined production tree. After #407 landed on
mainatb803c810, the tree-preserving post-land merge produced current head723f9a26; the inventory and unified verifier remain green. PR #409 was thensquash-merged into
mainate7e4a83f.Focused verification on the combined topology is green: float32 has 193 passed and 5
skipped; float64 has 198 passed; neither profile has failures or errors. The unified
static verifier passes, and the synchronized self-test rejects all 176/176 direct-flow
mutations.
Round-9 closure evidence at
1bff3e16:prek run --all-files: all 34 hooks passed without modifying files.Round 9 changes only the four certificate/test paths; no production
src/path changes.The closure state passes the strict audit gate with zero warnings or blockers.
A Windows portability follow-up at
3272af34makes the sealed-source hashesindependent of checkout newline convention. The certificate now canonicalizes UTF-8
text before hashing, so CRLF checkouts match the LF inventory while any non-newline
source change remains detectable. The expanded candidate-certificate battery passes
95/95 at both float32 and float64; inventory generation, the unified contract/policy
verifier, all 176 synchronized mutations, and an actual
core.autocrlf=trueclone alsopass. This follow-up changes three test/tooling paths and no production
src/path.Stack and related issues
This PR was stacked on
feat/continuous-outerand merged after #407. Thecollective, gated-edge, folded-shock, and transition-plan work touches the same
Bellman, transition, parameter, and simulation boundaries, so it is kept on the
single audited branch rather than split after the fact.
architecture above.