Add the endogenous-grid solver family: EGM, DC-EGM, and NEGM - #390
Conversation
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
Benchmark comparison (main → HEAD)Comparing
|
e0fecdd to
44e8522
Compare
Addresses the findings from the PR #390 code review. Correctness: - step_core: NaN-poison the carry marginal-utility row on envelope overflow, matching the value rows, so an overflowed period cannot feed the parent a finite-but-wrong Hermite slope past the NaN diagnostics. - numeric_inverse: require finite bracket marginals in the log_well_defined gate, so a +inf marginal utility (steep CRRA at a near-zero c_lower) fails loud instead of running the log path on a non-finite endpoint. - FUES upper envelope (#387): scan exhaustively by default (n_points_to_scan=None -> every candidate). A bounded window silently accepts dominated points when more than the window's worth of off-segment candidates interleave between two points of one segment. An explicit finite width remains available as a speed-vs-correctness opt-in. The F4 strict xfail becomes a passing test; a companion test pins the bounded-mode tradeoff. Cleanup and documentation: - regime_template: delete the dead value_transform/inverse_value_transform special-case orphaned by the certainty-equivalent merge. - validation: correct the module and helper docstrings that overclaimed the grid rules (batch_size is honored; only distributed is rejected) and that wrongly listed inverse_marginal_utility as required. - asset_row: docstring "weakly ascending" -> "strictly increasing" to match the enforced resources-monotonicity check. - continuation: document why a device-sharded carry into a DCEGM parent is unsupported, that three existing rules already make it unconstructible, and concrete ideas for lifting the restriction. A regression test pins the fence. Tests: ty, ruff, and the EGM/DCEGM/simulation suites pass (525 passed, 3 skipped); FUES parity suites green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0123hk3uzBgBeZown4eBBQ2C
mj023
left a comment
There was a problem hiding this comment.
Great additions, I was obviusly not able to check everything super thoroughly, but the algorithms look interesting and most of it seems to be well tested againt multiple benchmarks. Just some general thoughts:
General Structure:
I think the structure of the EGM folder is a bit cluttered, if possible, I think it would be better if the solvers could be split up so that each one has their own submodule.
Solvers:
As I understand it, the solvers are all specialized for slightly different models, which is fine for now. I feel like there is some overlap between the solvers, for example the one_asset_egm_step could reasonably be solved by the standard DC-EGM. I think we should overall try to find the most general solver that works well on GPU and focus on that for further development, maybe even remove some of the others later.
Upper Envelope Kernels:
I think here the same applies as for the solvers, we should try to at some point determine which methods are good and only support these. A good method would ideally return a constant number of gridpoints and be easily parallelizable to work on the GPU. The best fitting here will probably be the interpolation based algorithms like the query-algorithm or ltm-algorithm (which are nearly the same on the first look?). I am now quite sure, that both can also be implemented in a way, such that they are on average close to
Takes #390's convex per-cell hull ownership resolution ahead of the group-1 split moves. `segment_envelope.py` carries #390's 382-line rewrite (the 560-line version was identical on both sides before the push, so it merges without conflict), plus the new `cell_hull.py` and `double_double.py`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BNQkLG5QzXcgGdpgH6sN4u
Taken from `feat/dcegm` (PR #390, which is where this change actually lives -- PR #400's `pyproject.toml` is byte-identical to this branch's, so cherry-picking from #400 would have been a silent no-op). ONLY the aarch64 hunks are taken. #390's `pyproject.toml` diff also reorders the ruff `per-file-ignores` table and drops the `ARG001` ignores for `tests/test_certainty_equivalent.py` and `tests/test_temporal_aggregation.py`; both files exist here and use those fixtures, so importing that churn would red the lint for reasons unrelated to CUDA. `pixi.lock` is regenerated rather than patched, because adding a platform forces a re-solve. Checked, not assumed: - the aarch64 wheels are the right interpreter -- `jax_cuda13_plugin-0.11.0-cp314-cp314-manylinux_2_27_aarch64.whl` against this workspace's py3.14; - `nvidia_cudnn_cu13-9.24.0.43-...aarch64.whl` resolves. This is the one that matters: without cuDNN, jax loads the plugin, logs a line, and falls back to CPU roughly 20x slower -- a "GPU run" that silently is not one. Any run on that box should still assert `jax.default_backend() == "gpu"`; - no existing platform lost packages -- linux-64 2621 -> 2621, win-64 615 -> 615, osx-arm64 592 -> 592. The deleted lock lines are block reordering from inserting linux-aarch64, not removals; - `tests/solution/test_envelope_query.py` still passes locally under the regenerated lock (10 passed on the round-9 scale regressions). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #390 +/- ##
==========================================
+ Coverage 91.25% 91.42% +0.16%
==========================================
Files 177 249 +72
Lines 16956 24630 +7674
==========================================
+ Hits 15473 22517 +7044
- Misses 1483 2113 +630 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Brings across #400's `06bf038` (pin the envelope owner across eager, JIT and vmap) and `a60a092` (bound a certified margin by its own residual), #390's `af4b05b`/`ec96c12`/`ffb25e2`/`cdd4f65`, and the merges that carried them. One conflict, and it is NOT architecture: `_build_Q_and_F_per_period` in `src/_lcm/regime_building/processing.py` gained a keyword-only parameter on both sides -- `reachable_targets` from #390's declared-reachability work, and `period_to_regime_v_interp` from this branch. Both are needed and both are keyword-only, so they union; nothing is dropped. Per-file derivation with `git log 114b69b..18f5491 -- <file>` confirms no other conflicted file, and in particular `query.py` did not re-conflict: the ownership resolution from the previous merge holds. `tests/solution/test_envelope_margin_bound_is_outward.py` is REMOVED here, and this is the one judgement call in the merge. It arrives with `a60a092` and imports `_value_quotient` from #400's `query.py` -- an internal of the envelope architecture that the recorded three-branch resolution supersedes on this branch, so the symbol does not exist and `ty` fails. The property it asserts, that a certified margin's bound is outward and never understates its own error, is NOT superseded and applies to this kernel too; `certified_quotient_margin` is present here. Porting it against this branch's `_dd_div` is owed work, filed in the fleet coordination directory rather than dropped. Verification: no line added by either parent is absent from the result (the one both-touched file checked line-by-line; 11 theirs-only and 66 ours-only files byte-identical to their source parent); `prek run --all-files` clean, run by hand because hooks do not re-run on a merge commit; `pixi install --locked` clean for default and tests-cpu; `pixi lock` a no-op. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PVSLmeMo6iWUTcNER7bvM6
Brings the combined cascade hop to #409: bf61b14 (#400), 8f31991 (#406 transition-dependence fix), ddbec5b and 5b4339e are now all ancestors. One conflict, in `Q_and_F.py`, and both sides were right about different things. #407 `d9aec7d` transforms a stateless target's value BEFORE the expectation: `QuasiArithmeticMean` is `CE = g^-1(sum_r p_r * E_w[g(V'_r)])`, so `transform` precedes every expectation including the regime-transition one, and `inverse` applies once to the finished sum. Previously a stateless target was probability-weighted raw and only then inverted, leaving the transformed domain. #409 routes stateless targets through `mixture_terms` rather than a separate `E += p*V`, so they get `zero_safe_weighted_term` (a zero-probability +-inf is masked to exactly 0 instead of yielding NaN) and the value-ordered, alpha-renaming-invariant reduction. On this branch dissolution makes `-inf` continuations ordinary, so a raw product is a live NaN source, not a hypothetical. Resolved by doing both: CE-transform the stateless value, then seed `mixture_terms` with the transformed value. `_sum_regime_mixture` forms `sum_r p_r * g(V_r)` zero-safely and `ce.inverse` is applied once to the result, which is exactly `g^-1(sum_r p_r * g(V'_r))`. Also drops `test_collective_solve_default_call_still_returns_bare_mapping` (external test-suite dedup audit, round 1): it is a local mirror of `test_collective_regime_simulate.py::test_public_model_solve_default_return_shape_is_byte_identical`, which is the module it already imports its model helper from, and whose assertions are a strict superset. The file's own singleton guard, `test_singleton_solve_bare_call_returns_legacy_mapping_not_a_tuple`, is retained -- that is the invariant this module exists for. The other five deletions from that audit are shared with `main` and went to #390 as a patch. Verified: ty clean; tests/regime_building + the two touched test files 379 passed; the certainty-equivalent bearing files (test_temporal_aggregation, test_dcegm_validation, test_nbegm_epstein_zin_composite_flow, test_simulate_terminal_rows) 43 passed. KNOWN COVERAGE GAP, not closed here: no test appears to construct a model with BOTH a certainty equivalent AND a stateless target, so the intersection this resolution creates is exercised by neither side's tests. `d9aec7d` itself shipped without tests. The two halves are covered separately; their combination is not.
Carries #390's `45f1abe` (enter a target-only IID process at its own unconditional law) plus nine other nb-egm/dcegm commits. Two conflict hunks in processing.py, both unions. The second collides #390's restructured process-synthesis block with this branch's model-wide `_validate_all_conditioned_processes` sweep; took the restructure and kept the sweep ahead of it, since it validates declarations off `all_grids` and does not depend on the target scoping. One clean-merge break, no conflict: #390's carried-process call site is `_get_weights_func_for_process(name=process, grid=grid)`, and `grids` survived only as an optional `MappingProxyType({})` default. A state-conditioned process resolves `state_conditioned.on` against that mapping, so an empty one fails the build with a misleading "must name a DiscreteGrid". Threaded `grids=all_grids[user_regime]` back in. Flagged upstream, not fixed here: `_get_entry_weights_for_process(*, name, grid)` has no `state_conditioned` handling, so a TARGET-ONLY conditioned process would be entered at the unconditional row and silently ignore its conditioner. Nothing currently forbids that declaration. prek --all-files green (ruff, ruff-format, ty); 290 tests pass across tests/regime_building/, test_state_conditioned_{model,shocks}, test_processes, test_runtime_process_params and the three touched solution modules. Not the full suite.
…ansitions. Clean merge, no conflicts. Carries #390's target-only IID process entry (`45f1abe`) and the state-conditioned `grids` repair from the step below. prek --all-files green (ruff, ruff-format, ty); 290 tests pass across tests/regime_building/, test_Q_and_F, test_coarse_law_provenance, test_processes and the three touched solution modules. Not the full suite.
Clean merge, no conflicts. Carries #390's target-only IID process entry (`45f1abe`) through #405/#406. `query.py` unchanged, blob e990fb6 -- the object under Pro audit. prek --all-files green (ruff, ruff-format, ty); 244 tests pass across tests/regime_building/, test_solvers, test_outer_search, egm/test_branch_aggregation, test_continuous_outer_audit_regressions and the two touched solution modules. Not the full suite.
An age-specialized constraint that stopped being resolved per period would still solve. The model builds, nothing raises, and the only trace is a value function restricted by another period's rule — so the coverage that matters is whichever assertion notices that, and the file had none. Its age-specialized constraint is deliberately slack, so holding every period to one closure changes nothing it asserts. The last active period is the only period a uniform-cap reference can attribute to an age: its continuation is terminal, so its value depends on no other period's cap. Solving that period under a uniform cap fixed at the same age must reproduce the specialized solve exactly there, which fails under a freeze and under a permutation alike. The attribution reaches no further, and the docstring says so rather than implying a coverage the arithmetic cannot give. The second test is what keeps the first from being empty. A reference every age would have produced makes an equality against it say nothing, and that is not hypothetical: the tightest cap admits only the lowest action node, so with log utility its value function is identically zero, and an assertion written against that reference reduces to a claim about the other solve alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
…allable A constraint reached a solver as a bare predicate wearing an attribute that said what it had been declared as. Reading that attribute is textually quiet and type-clean, so a solver that stopped setting it kept compiling and simply stopped proving anything — the verdict flipped with nothing to see. Normalization now produces a `ProcessedConstraint` for every declaration, carrying both what the constraint says as an inspectable condition and the object the user wrote. A solver asks the condition for its boundary surfaces and proves a lower bound against its own savings grid by the name the bound is on, so a comparison written by hand on some other name is an ordinary constraint rather than a silently mis-keyed one. An opaque constraint is passed through as the callable its author wrote rather than rebuilt from the surrounding function pool: it already carries annotations, and the DAG requires every consumer of a name to annotate it as its producer does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
The seeded wealth sat outside both the wealth grid and the consumption grid, so the budget precondition the test asserts before simulating was comparing a value the model could never hold. Seed inside the grids and assert the containment, so the precondition tests the budget rather than the seeding. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
A bound that has to be looked up is named by a reference to whatever computes it rather than by an indexing node, so the tree stays a statement about names. Record the consequence alongside the rule: a solver cannot read such a bound's number off the tree, and turning a surface plus a name into something it can prove against its own grid is the boundary compiler's job. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
A grid holds its lowest node at the active float dtype, so reading that node back as a Python float produced a different representation of the decimal the user had written. Declaring the very number a grid was built from was then refused for every decimal the dtype cannot hold exactly — at float32, which is what runs outside the test suite, that is most of them. The comparison now happens in the grid's dtype, where it is exact. The check also reached the node through `to_jax()`, which for a grid supplying its points with the params raised an error naming neither the grid nor the regime, ahead of the rule that refuses such a grid properly. That refusal moves into a helper both sites call, so the diagnosis no longer depends on which check runs first — and a solver with no grid rule of its own gets it too. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
The engine drops a declared post-decision lower bound wherever a solver enforces one through a savings grid, and it decides that from the solver's family and the regime's margin — which covers NEGM. Only DCEGM ran the check that compares the declared number with the grid, so a NEGM regime could state a limit its grid does not carry, build without complaint, and lose the constraint altogether: the declaration was discarded as one the grid guarantees, with nothing having compared the two. The nested solve inverts on the inner solver's grid, so that grid's lowest node is the limit a NEGM regime enforces and the one the simulate-phase mask is built from. NEGM now proves a declaration against it, and a disagreement names both numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
A solver does not have one place where it meets a constraint. A nesting solve produces candidates through more than one pipeline in a single phase, and rewrites its function pool between them, so `(constraint, regime, phase)` is too coarse a unit to carry a verdict: it has to drop one of the pipelines, and a constraint that reaches no verdict is neither honoured nor refused. That one has no symptom — it surfaces as a wrong policy rather than as an error. A route is one solver's ordered candidate-production pipeline for one phase, and a site is one point along it. The planner walks a route's sites in order and gives every constraint exactly one terminal disposition on every route, which lets the same borrowing limit be enforced by construction where a savings grid enforces it and evaluated in simulation, where no such grid exists. What a constraint needs is resolved through the site's own pool rather than read off its surface, so `spendable >= 0` and the same requirement spelled over `spendable`'s own leaves are disposed of alike. Reading the surface would make one wording solvable and the other refused. Evaluating at an earlier site beats proving at a later one: a redundant predicate costs time, while treating a constraint as discharged by a construction that is never reached is silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
A site's allow-list now carries the whole statement about evaluation: `None` is no restriction, an empty frozenset says the site evaluates nothing, and a non-empty one is the allow-list a constraint's leaves must be a subset of. Two things went with it. A site no longer declares which names it binds separately from which it can read, because two statements of one fact agree until they do not and nothing observable says which rotted — a branch site that binds its outer node per candidate and launches an inner kernel calling no predicate would have had one field saying "in scope" and the other "evaluates nothing". The cost is that a site which binds a name and does evaluate must list that name itself; forgetting it refuses the constraint at model build with the missing name in the message, which is loud, where the other way round the site would claim a constraint nothing there evaluates. And evaluating nothing is a rule rather than a consequence of the subset test. The subset test gets that case wrong in one corner: a constraint needing no name at all is trivially within an empty allow-list, so a condition comparing two literals — or a zero-argument callable — was handed to a kernel that calls no predicate. Degenerate to write, perfectly declarable, and silent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
Grid search searches the whole state-action product, so every name a constraint could read is bound where it evaluates and its route is one unrestricted site. Plain EGM's kernel calls no user predicate at all: its solve route is a single site that evaluates nothing, which is a place a constraint can still be discharged and never one where it can be called. DC-EGM builds its feasibility predicate once per discrete combination, before the continuous inner stage, so what is readable there is every discrete state, every passive continuous state, every discrete action and the regime's params — and not the Euler state or the continuous action, which the inversion produces rather than binds. Simulation is a different pipeline with a different answer for all three: the subject's realized action is in hand, so the check runs over a whole candidate. That a constraint refused on the solve route is evaluated on the simulate one is the thing a single verdict per regime could not say. `build_constraint_routes` returns `None` by default rather than being abstract or defaulting to an unrestricted route. Abstract would turn every existing custom solver into a construction-time error naming a method its author has never heard of; a permissive default would claim, of every solver nobody has written down, the opposite of the truth for any that evaluates nothing. A route's key carries `period_group=None` when the route is the same in every period. One route per period group from a solver that resolves its pool alike at every age would put an entry per group in the plan where there is a single fact, and a coverage count over those pairs would read a constant as evidence of grouping. So the context hands over the active periods rather than a partition, and a solver that genuinely rebuilds per group partitions them itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
Simulation is the phase's pipeline, not a solver's. It walks the regime's DAG on each subject's realized states and realized action, so a whole candidate is in hand and every name is computable — true of every solver shipped, whatever it does when solving. Spelling that out separately in each producer would put six copies of one declaration in the tree, agreeing by convention until one did not, and the disagreement would sit in a field nobody compared. A solver whose simulate phase genuinely differs writes its own route rather than calling this, which then reads as the deliberate departure it is. The route builder takes the proofs the site consults, so an endogenous-grid family can pass the one its savings grid's lowest node establishes: the simulate-phase mask is built from that node, which is what makes the bound already true there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
…ishes A savings grid's lowest node is the borrowing limit the solve enforces, and the simulate phase enforces the same number through the mask synthesized from it. So a declared bound on the post-decision state that grid spans is discharged by construction on both routes, and the solver says so through a proof its sites consult rather than through the constraint being absent. The declaration stays a claim. It is checked against the grid when the model is built, which is what makes proving it different from ignoring it, and only a bound on that one state is proved — a lower bound on any other name says nothing about the grid, so nothing has proved it. `ConstraintRoute.sites` says which reading of a site is meant. A solver's candidate production may have any number of stages, and only those at which a constraint can be evaluated, proved, or compiled belong there: a stage that can decide none of the three cannot change a plan, so declaring it would put pipeline shape into the field the ledger's decisions are counted over. Which is why an endogenous-grid solve declares one site and not one per stage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
…empty one The witness for "a constraint that reads nothing is refused where nothing is evaluated" is only a witness if it really reads nothing, and the walk that decides that resolves a constraint through the site's own pool rather than reading its surface — so the two assertions are different claims and the test needs the one the planner consults. It also needs a witness that comes back non-empty in the same run. With a small pool a broken walk returning nothing is indistinguishable from correct behaviour on every input, so an empty answer on its own is a negative result from an instrument never shown firing. Both read the pool off the site rather than naming it again. A second reference to the same object asks a different question that happens to share an answer, which holds until a producer changes which pool it hands over. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019cJmPwbuLApd35wGJWWMt5
A constraint's fate is now settled where the model is built, by planning it over the routes its solver declares for that phase and handing the kernel only the ones a route evaluates. One a route proves by construction is discharged rather than dropped: the difference is that the ledger holds a reason for it, so a constraint can no longer leave the set without one. The refusal changes with it. A constraint DC-EGM cannot read is refused by the route that could not meet it, naming the pipeline and what the constraint still needed there, in place of a rule about continuous variables that named neither. The same declaration is accepted under grid search, which is the point: which constraints a regime may state is a fact about the pipeline, never about the constraint. The simulate phase synthesizes its own budget mask for the endogenous-grid family, after the declared constraints have been narrowed. Injecting it downstream of the plan left one constraint reaching a kernel with no disposition -- the single state the ledger exists to exclude -- so it is planned too, against the pool it is evaluated over. Two things stay that the route model will eventually replace, both because removing them would lose a refusal rather than a redundancy: - The savings-grid drop, on the no-plan path only. `build_constraint_routes` returning `None` says a solver's pipeline has not been written down, not that it is unrestricted, so nothing is planned and today's behaviour holds. That is NEGM, whose declared bound resolves through to the Euler state its inner combination pool does not bind; removing the drop before its routes exist turns a silent drop into a failure at solve time. - The per-combination argument check in the kernel's scope pass, which runs on the age-resolved pool where the plan runs on the declared one. The two allow-lists are the same set only if a solver's passive states are exactly its non-Euler continuous states, which is unmeasured. A phase builds one constraint set, so two routes of one phase that disagree about a constraint have no single answer to hand it. Every solver declaring routes today declares one per phase and cannot disagree; a nesting solver whose branches evaluate different sets is refused until the kernel takes one set per branch, rather than being served whichever route was asked last.
Give the adjuster, keeper, and simulation pipelines explicit constraint routes, including their rewritten function pools and period groups. Resolve age-specialized functions before the route ledger walks their annotations, and remove the legacy proof-by-deletion fallback now that shipped solver routes own discharge.
Require construction-time nodes for NEGM's outer state, inner savings, and outer search grids. These grids determine continuation layout, inversion, and outer candidates and therefore cannot defer their points to runtime parameters.
Record the repository convention for delegating test batteries, preserving verbose nodeids, and reporting exact results from JUnit XML.
timmens
left a comment
There was a problem hiding this comment.
Very nice!
-
Plain EGM bypasses its lower-bound proof
(src/_lcm/solution/egm.py):validate_modelrejects every constraint before
route planning, including a matchingpost_decision_lower_bound. Let
proof-carrying constraints reach the planner and reject only unmet ones. -
Conditionis broader than current use (src/lcm/condition.py,
src/_lcm/constraints/): the only shipped structural consumer recognizes one
numeric lower bound; no boundary compiler exists. A typedLowerBoundor
InequalityConstraintwould cover this with much less complexity. What concrete
second consumer justifies the Boolean DSL? -
name_in_dagleaks internals
(src/lcm/consumption_savings_regime.py): rename it tooutputorname, or
infer it. -
Regime.solveroverstates support (src/lcm/regime.py): it lists EGM,
DCEGM, and NEGM, although these require specialized regime classes. Document the
pairing explicitly.
|
Thanks! 1, 3, 4 are implmented.
The genesis here is that we needed it in #400, but then introduced it here so we don't introduce a transient interface. #400 already contains the substantive boundary-processing algorithm: it maps declared case and piecewise-affine metadata into asset-axis breakpoints and interval-specific EGM problems. But your comment was constructive in the sense that some elements have not been wired in yet, will do so now! |
Summary
This PR adds the endogenous-grid solver family to pylcm and makes solver choice an
explicit part of a regime's economic structure:
EGMfor the classic one-dimensional consumption-savings problem;DCEGMfor a liquid Euler margin combined with discrete choices or othernon-Euler axes;
NEGMfor a liquid inner DC-EGM problem nested inside an outer search over adurable or illiquid margin;
approximate alternatives.
The implementation also introduces a constraint-routing contract. A solver must
account for every applicable constraint on every candidate-production route: it
must evaluate it, prove it by construction, compile its boundary, or reject the
model during construction.
Motivation
Grid search spends most of its work evaluating a dense continuous-action grid.
EGM reverses the Euler equation instead: it starts from an exogenous
post-decision-state grid and constructs the current-state grid on which the Euler
condition holds. This makes finely resolved consumption-savings problems much
cheaper, but only when the model exposes the economic roles and dependency
structure that the algorithm needs.
DC-EGM extends that idea to models with discrete choices. Those choices generate
overlapping, potentially non-concave value branches, so the solver must take an
upper envelope. NEGM adds a second continuous margin whose kinks or adjustment
costs make a second Euler inversion inappropriate; it retains DC-EGM for the
liquid margin and searches the outer margin directly.
Public interface
Solver and envelope names live under
lcm.solvers. The specialized economicdeclarations live under
lcm.consumption_savings_regime:When resources are simply the liquid state, the state can fill that role
directly; an identity
resources(wealth) -> wealthDAG node is unnecessary.LiquidMargin.resourcescan instead name a resources function or aNetOfAdjustmentCostdeclaration when the budget has another investmentdecision with adjustment cost (e.g., housing).
The regime declares economic roles; the solver object contains numerical choices.
This keeps names such as
wealth,consumption, andsavingsout of solverconfiguration and lets validation explain structural mismatches at
Model(...)construction.
Regime/GridSearchConsumptionSavingsRegime/EGMConsumptionSavingsRegime/DCEGMNestedConsumptionSavingsRegime/NEGMConsumptionSavingsRegimeandNestedConsumptionSavingsRegimealso acceptGridSearch, so the same economic declaration can be used as a reference solution.A plain
Regimecannot select an endogenous-grid solver because it does not declarethe required margin roles.
Constraints and
ConditionOrdinary constraint callables remain fully supported:
For constraints whose comparison structure matters to a solver, the same Boolean
can be written as a
Condition:Conditionsupports named/literal and named/named comparisons, intersections(
&), unions (|), complements (~), and implication. A reference can name astate, action, DAG output, parameter, or declared margin role.
The callable and structured forms have the same meaning and produce the same
feasibility mask. Use
Conditionwhen:from its grid or compile a boundary; or
It is not otherwise required.
Condition.from_callable(...)preserves an arbitrarycallable but is intentionally opaque: it does not give a solver comparison structure
that was absent from the callable.
This distinction matters for endogenous-grid solvers because they do not construct
the same candidates as grid search. For each solver route and phase, every constraint
receives exactly one disposition:
shape.
The borrowing limit is the main example. The lower node of an EGM-family savings
grid already enforces a lower bound on the post-decision liquid state. Authors can
make that contract explicit with either a structured comparison or the margin-aware
helper:
Under
EGMandDCEGM, pylcm checks that the declared value equals the savingsgrid's lower node and then proves the constraint by construction. Under
GridSearch, the returned ordinary predicate is evaluated normally. The declarationis optional if the savings grid is intended to be the sole source of the borrowing
limit. An opaque callable that happens to spell the same inequality cannot be silently
treated as proven, because the solver cannot inspect what it says.
Solver contracts
The endogenous-grid solvers differentiate the model's declared utility and law of
motion rather than reconstructing a hard-coded savings equation. An analytical
inverse marginal utility function is optional; when absent, pylcm differentiates
utility and numerically inverts marginal utility.
Model construction validates the contract before solving. Among other checks, it
verifies:
nodes;
needs them;
backend.
Unsupported structures therefore fail during
Model(...)construction instead ofbeing accepted and producing a plausible but invalid policy.
Upper envelopes
DCEGM.envelopeis a typed configuration:ExactEnvelopecertifies ownership among the finite candidates supplied to it;FUESEnvelopeprovides the Fast Upper-Envelope Scan;RFCEnvelopeprovides Rooftop Cut;LTMEnvelopeprovides a quadratic local-upper-bound reference;MSSEnvelopeprovides a HARK-style left-to-right segment sweep.ExactEnvelopeis the default. It relies on pylcm's own compiled exact-affine kernel,built and shipped as part of the package, to make candidate-ownership decisions in
exact integer arithmetic over the stored floating-point operand bits. If that kernel
is unavailable, selecting
ExactEnveloperaises a capability error rather thanfalling back to floating arithmetic that cannot provide the same guarantee. Builds
that deliberately do not need the certified backend can set
LCM_SKIP_EXACT_AFFINE=1and select one of the approximate envelope configurations.The certification is deliberately scoped: it identifies the envelope of the finite
candidate set presented to the kernel. The resolution of a sampled constrained branch
remains controlled by
DCEGM.n_constrained_points.What this changes internally
general
Regimemodule.interpolation, and policy publication.
parameters, and target-specific laws through DC-EGM's candidate construction.
compilation.
constraint-routing, envelope, installation, and import-surface tests.
Documentation
The documentation starts from model and constraint shape, then explains which solver
fits that shape. It includes:
EGM/DCEGM/NEGMdecision boundary;Conditionauthoring;savings grid;