diff --git a/system/decisions.md b/system/decisions.md index 991841d..583104a 100644 --- a/system/decisions.md +++ b/system/decisions.md @@ -305,9 +305,9 @@ removes it here. - **Association-set identity.** The slot-identity table on [products](products) is the sole statement. An association set's identity adds the field, crossmatch settings hash, sorted source-set - identities, and a hash of its base's own identity. Withdraw the - proposal's shorter wording, which omitted the base, and its table - copy. (2026-09-27) + identities, and a hash of its base's own identity; the proposal's shorter + wording, which omitted the base, is withdrawn with the table copy it + sat in. (2026-09-27) (decision-pending-sharing-rule)= - **Sharing rule widening.** The specification lets any run read diff --git a/system/runs.md b/system/runs.md index ecf8416..f02b257 100644 --- a/system/runs.md +++ b/system/runs.md @@ -2,22 +2,23 @@ **Status: DRAFT** -Companion to the specification's Runs section, the [stage contract](stage-contract) -and the [products](products) page: how runs, units of work, attempts, product -instances, result sets and the three custody states are held in the database -and laid out in storage. +Runs, units of work, attempts, product instances, result sets and the +three custody states have database records and storage locations. This +page defines them, alongside the specification's Runs section, the +[stage contract](stage-contract) and the [products](products) page. + ## In plain terms -A run is a row that says who started a pass over which inputs with -which code and settings, and whether its outputs are a person's or the -project's. Each piece of work in the run is a unit; each try at a unit -is an attempt with its own output folder. Every product a finished -attempt made, file or database rows, is an instance row. Three states -on the instance row say whose it is and whether consumers see it. -Promotion changes those rows under one lock and writes down what it -changed, so it can be undone. Files never move; only rows change. -Deleting a scratch run is one guarded operation, and nothing else -deletes run data. +A run records who started a pass over which inputs, with which code +and settings, and whether the outputs belong to a person or the +project. Each piece of work is a unit; each try is an attempt with its +own output folder. Every product a finished attempt makes, whether a +file or database rows, has an instance row. Its custody state says +whose it is and whether consumers see it. + +Promotion changes those rows under one lock and records the changes +so they can be undone. Files never move. Scratch-run deletion is one +guarded operation; nothing else deletes run data. ## Tables @@ -44,478 +45,539 @@ uniqueness constraint would block two runs holding the same logical product it is widened to include the run or set. Nothing is renamed and nothing is dropped. -## Rules - -**Runs.** A run's kind is fixed at creation and decides the custody of -everything it makes: a scratch run makes scratch, a production run makes -candidates. A run freezes its input selection at creation. A run may -bind a complete current result-set instance from another run as a -frozen input; that binding stays valid if the instance is later -superseded, because the dependency retains it. Scratch result sets are -usable only within their own run. Production runs use the one -production database; scratch runs may name a trial database. A run -becomes eligible to finish once every unit is terminal, but finishing -is an explicit `finish` and not automatic; a finished run is never -reopened. Seeding a new run copies -configuration and permitted input selections; it does not authorise -reuse of another run's scratch outputs. A run the processing-date loop -creates carries its spec's owner as its own owner and the spec's -location as its `input_selection_ref`; a promotion the loop performs -records `who = scheduler` (the [loop](loop) page). - -`run create --seed --only-failed` is the recovery form of -seeding, and what it copies -depends on the seed's own kind, since a scratch run's outputs are -usable only within their own run and a production run's are not. A -production seed's recovery run copies the seed's release (or code -revision and image digest), settings and input refs, lane, resource -profile, database target, max attempts and check-policy ref, sets -`purpose` to name the seed it recovers, and takes `selected_stages` -from the seed's own list starting at the earliest stage holding a -non-complete unit: one in state `failed` or `cancelled`, or `running` -with a job-less or `lost` attempt. From that position it creates one -new `pending` unit for every non-complete unit anywhere in the seed, at -that stage or any later one, each carrying `units.seeded_from_unit` and -the seed unit's own frozen `unit_inputs` bindings copied across: a -copied binding points at the same producer instance the seed unit -already depended on, so the existing deletion fence already protects -it, the same as any other frozen input binding of unfinished work. -`run start` skips a stage the seed already completed and walks straight -to the seeded units. A scratch seed's recovery run instead recreates -every unit from the first stage, carrying over only the first stage's -inputs and settings from the seed, since later stages must regenerate -their inputs within the new run exactly as a fresh run does. It -refuses, exit 64, when the seed -has no non-complete unit or is deleting or deleted; `--only-failed` -without `--seed` is a usage error. +## Runs + +A run's kind and input selection are fixed at creation. Scratch runs +make scratch outputs; production runs make candidates. Production uses +the one production database, while scratch runs may name a trial +database. A run may bind another run's complete current result-set +instance as a frozen input. The dependency retains that instance, so +the binding stays valid after supersession. Scratch result sets are +usable only within their own run. + +A run is eligible to finish once every unit is terminal. Finishing +requires an explicit `finish`; a finished run is never reopened. +Seeding copies configuration and permitted input selections without +authorising reuse of another run's scratch outputs. A run created by +the processing-date loop takes its owner from the spec and records the +spec's location as its `input_selection_ref`. A loop promotion records +`who = scheduler` (the [loop](loop) page). + +### Recovery from a seed + +`run create --seed --only-failed` creates a recovery run. What it +copies depends on the seed's kind. For a production seed, it copies +the release (or code revision and image digest), settings and input +refs, lane, resource profile, database target, max attempts and +check-policy ref. It sets `purpose` to name the seed it recovers and +takes `selected_stages` from the seed's list, starting at the earliest +stage with a non-complete unit: `failed`, `cancelled`, or `running` +with a job-less or `lost` attempt. + +At that stage and every later one, each non-complete seed unit becomes +a new `pending` unit carrying `units.seeded_from_unit` and a copy of +its frozen `unit_inputs`. Each copied binding names the same producer +instance, protected by the existing deletion fence like any frozen +input of unfinished work. `run start` skips stages the seed completed +and goes straight to the seeded units. + +A scratch seed's recovery run recreates every unit from the first +stage. It carries over only that stage's inputs and settings; later +stages must regenerate their inputs within the new run, as in a fresh +run. Recovery refuses with exit 64 if the seed has no non-complete +unit or is deleting or deleted. `--only-failed` without `--seed` is a +usage error. + +### Expiry Scratch runs receive a default `expires_at` of fourteen days after -creation. `pinned` holds a run past that date. An expiry sweep deletes -unpinned scratch runs past their `expires_at`, through the same -operation as `run delete`. The sweep runs as an explicitly authorised -actor, not the run's owner; under the run lock it re-checks the run's -kind, pin and expiry and refuses the delete if any no longer holds. -The sweep's warning mechanics are not -decided here. - -**Units.** A unit is pending until its declared input set is complete -and its upstream attempts are selected. It then becomes ready, and -running when an attempt is allocated. Selection of a successful -completed attempt makes it complete. A retryable failure returns it to -ready while attempts remain; otherwise it becomes failed. Cancellation, -or a permanently failed required dependency, makes an unstarted unit -cancelled with a recorded reason. Complete, failed and cancelled are -terminal. Unit state is stored, not derived. Inputs are bound in -`unit_inputs` before execution and retained for retries. `register`'s -own unit id is `/` (for example -`admit/e20260821001234/SCA07`, `difference/e20260821001234/SCA07`), -derived from the manifest `register` reads: one `register` unit follows -every producer, with no occurrence counter, and it is runnable singly -for one producer at a time. Rows written under an earlier unit-id form -stay as they are. A unit created -by `run create --seed --only-failed` carries -`units.seeded_from_unit`, the seed run's unit it re-runs; it enters the -state machine above as an ordinary new `pending` unit, with no separate -path. - -**Input-set composition.** A stage declares its input set by product -kind and role. The launcher resolves each entry from this run's own -instance of that kind for the unit first; failing that, from the run's -frozen input selection, using `dev`'s selection rules (the current -reference by field and filter, the current PSFs by filter and -detector). It never resolves an entry from an unrelated run unless the -input selection names that instance explicitly. The resolved binding is -written to `unit_inputs` before execution, and the manifest a stage -reads is generated from that binding; a hand-composed manifest is a -test path, not how a production run assembles its inputs. Before any write for a submission (`run submit`, `run -start`, `run local`), the launcher reads `manifest.json` at the -`--inputs` location, local or `s3://`; collects every output entry's -`instance` and every `inputs.result_sets` entry -- not `inputs.products`, -which are what the *upstream* attempt read, not this unit's own -binding; creates the unit; binds, through `bind_unit_inputs`, the -collected names that are registered product instances; commits both -together; and only then allocates the attempt. A name that is not a -registered instance -- a delivery manifest, a dev-era template entry -- -binds nothing and is logged, not refused. A manifest that is absent, -invalid or unreadable for a non-network reason refuses the submission, -exit 65, before anything is written; a network error refuses it with -exit 75 instead. Binding is idempotent per (unit, instance): a retry -and a seeded `--only-failed` re-run both re-read the manifest and bind -nothing new. Binding takes each producer instance's run row for share -and refuses, exit 65, when that run is deleting or deleted, as -`register_manifest` does; an uncommitted binding therefore holds the -producer's run against deletion. The fence also covers `run inputs` and -the loop's own binding. The deletion guard's basis is this binding. - -One primitive backs both composers, `run inputs` and `run start`'s -`compose_inputs` and the loop's maintain, crossmatch and alerts sites: +creation. `pinned` holds a run past that date. The expiry sweep deletes +unpinned scratch runs past their `expires_at` through the same +operation as `run delete`. It acts under explicit authorisation, +rather than as the run's owner. Under the run lock, it re-checks kind, +pin and expiry and refuses deletion if any condition no longer holds. +The warning mechanics are not decided here. + +## Units + +A unit is pending until its declared input set is complete and its +upstream attempts are selected. It then becomes ready, and running +when an attempt is allocated. Selecting a successful completed attempt +makes it complete. A retryable failure returns it to ready while +attempts remain; otherwise it becomes failed. Cancellation or a +permanently failed required dependency cancels an unstarted unit with +a recorded reason. Complete, failed and cancelled are terminal. Unit +state is stored, not derived. Inputs are bound in `unit_inputs` before +execution and retained for retries. + +`register`'s unit id is `/`, for +example `admit/e20260821001234/SCA07` or +`difference/e20260821001234/SCA07`, derived from the manifest +`register` reads. One `register` unit follows every producer, with no +occurrence counter, and is runnable singly for one producer at a time. +Rows written under an earlier unit-id form stay unchanged. + +A unit created by `run create --seed --only-failed` carries +`units.seeded_from_unit`, identifying the seed unit it re-runs. It +enters the same state machine as an ordinary new `pending` unit. + +## Input binding and composition + +A stage declares its input set by product kind and role. The launcher +resolves each entry from this run's own instance of that kind for the +unit first, then from the run's frozen input selection. It uses +`dev`'s selection rules: the current reference by field and filter, +and the current PSFs by filter and detector. It never resolves an entry +from an unrelated run unless the selection explicitly names that +instance. + +The binding is written to `unit_inputs` before execution, and the +stage's input manifest is generated from it. Hand-composed manifests +are a test path; production runs assemble inputs through the launcher. + +### Submission + +Before any submission write (`run submit`, `run start`, `run local`), +the launcher reads `manifest.json` at the local or `s3://` `--inputs` +location. It collects every output entry's `instance` and every +`inputs.result_sets` entry. It excludes `inputs.products`, which +names what the upstream attempt read rather than this unit's inputs. +It then creates the unit, binds the registered product instances +through `bind_unit_inputs`, commits both together and allocates the +attempt. + +Unregistered names, such as a delivery manifest or dev-era template +entry, bind nothing and are logged rather than refused. An absent, +invalid or unreadable manifest refuses submission before any write: +exit 75 for a network error, exit 65 otherwise. Binding is idempotent +per (unit, instance): a retry and a seeded `--only-failed` re-run both +re-read the manifest and bind nothing new. + +Binding takes each producer instance's run row for share and refuses +with exit 65 if that run is deleting or deleted, as +`register_manifest` does. An uncommitted binding therefore holds the +producer's run against deletion. The fence also covers `run inputs` +and the loop's binding. These bindings are the deletion guard's basis. + +### Composition + +One primitive backs `run inputs`, `run start`'s `compose_inputs` and +the loop's maintain, crossmatch and alerts composers: `rapidpipe.runs.binding.bind_input_set(conn, storage, *, run_id, stage, -unit_kind, unit_id, dest, compose, reuse_existing=True)`. The -submission-time bind above (`run submit`, `run start`, `run local`, -through `bind_registered_inputs`) stays its own path, sharing with the -primitive only the id rule. Its order is fixed: check whether -`manifest.json` already exists at `dest` (refused when reuse is not -allowed -- `run inputs`, exit 64 -- and reused under `run start` and by -the loop); admit the unit, through `add_unit`, the run fence, before -anything is copied. Admission is the first write-side check of -every composition, before the template and producer manifests are read -and before an existing manifest is parsed, so a finished, deleting or -deleted run is refused, exit 64, before its inputs are inspected. -The composer then reads the existing manifest, or calls its own -`compose` callable, which does the copying and builds the manifest; -collects the manifest's instance ids by one rule, `manifest_instances`, -every output entry's `instance` plus every `inputs.result_sets` entry; -binds the registered ones through `bind_registered_inputs`; commits; -and writes the manifest to storage last, only when it did not already -exist. Reusing an existing manifest is therefore an idempotent re-bind -of that manifest's ids to the consuming unit, committed like any other -bind, never a skipped bind and never a rewrite. A manifest on storage -means its bindings were committed; if the write fails after the commit, -the next call finds no manifest and composes again. Both -`compose_inputs` and the loop's composer call this primitive, and a -test fails if either binds or writes a manifest on its own. So `run -inputs` binds a template's `inputs.result_sets` as well as its output -entries, and a deleting or deleted producer met at `run inputs`'s bind -exits 65. - -**Attempts.** Each try is an attempt with a fresh id and an exclusive -output location; an attempt has no disposition while queued or running. -Success requires a valid completion manifest and complete declared -outputs; exit zero alone is not success. Selection locks the unit row -and atomically sets its selected attempt and complete state; the -selected attempt belongs to that unit and the selection never changes. -Late results cannot reopen a terminal unit or replace its selected -attempt. The run records the maximum attempts per unit, counting the -first attempt and every retry. The launcher retries, never Batch: a -Batch job runs its container once, and a unit whose attempt ended -`transient` returns to ready and is submitted again as a new attempt. -Only code 75 and the approved infrastructure failures listed on the -[stage contract](stage-contract), "Exit codes", are recorded -`transient`; an exhausted allowance fails the unit. `lost` means the scheduler lost the job: unresolved -execution, treated as a possible writer until resolved. `submit_unit` -records the input-set and settings locations it resolved for the -attempt on `attempts.inputs_location` and `attempts.settings_location`, -alongside the frozen bindings `unit_inputs` already carries. `run reconcile --resolve-jobless [--older-than +unit_kind, unit_id, dest, compose, reuse_existing=True)`. Submission-time +binding (`run submit`, `run start`, `run local`, through +`bind_registered_inputs`) remains a separate path, sharing only the id +rule with the primitive. + +The primitive first checks whether `manifest.json` exists at `dest`. +`run inputs` refuses reuse with exit 64; `run start` and the loop reuse +it. Next, `add_unit` admits the unit through the run fence. This is the +first write-side check, before copying anything, reading template or +producer manifests, or parsing an existing manifest. A finished, +deleting or deleted run is therefore refused with exit 64 before its +inputs are inspected. + +The composer reads the existing manifest or calls its `compose` +callable to copy inputs and build one. The `manifest_instances` rule +collects every output entry's `instance` plus every +`inputs.result_sets` entry. The composer binds registered instances +through `bind_registered_inputs`, commits, then writes the manifest to +storage only if it did not already exist. Reuse commits an idempotent +re-bind of the existing manifest's ids to the consuming unit without +rewriting the manifest. A stored manifest therefore means its bindings +were committed. If the write fails after commit, the next call finds +no manifest and composes again. + +Both `compose_inputs` and the loop's composer call this primitive; a +test fails if either binds or writes a manifest independently. Thus +`run inputs` binds a template's `inputs.result_sets` and output +entries, and exits 65 if a producer is deleting or deleted at binding. + +## Attempts and retries + +Each try has a fresh attempt id and an exclusive output location. An +attempt has no disposition while queued or running. Success requires +a valid completion manifest and complete declared outputs; exit zero +alone is insufficient. Selection locks the unit row and atomically +sets its selected attempt and complete state. The selected attempt +belongs to that unit and never changes. Late results cannot reopen a +terminal unit or replace its selection. + +The run records a maximum attempts per unit, counting the first try +and every retry. The launcher retries; a Batch job runs its container +once. An attempt ending `transient` returns the unit to ready for +submission as a new attempt. Only code 75 and the approved +infrastructure failures listed in the [stage contract](stage-contract), +"Exit codes", are recorded `transient`. Exhausting the allowance fails +the unit. `lost` means the scheduler lost the job: unresolved +execution, treated as a possible writer until resolved. + +`submit_unit` records the resolved input-set and settings locations on +`attempts.inputs_location` and `attempts.settings_location`, alongside +the frozen bindings in `unit_inputs`. + +### Job-less attempts + +`run reconcile --resolve-jobless [--older-than SECONDS]`, default 600 seconds, first looks for a scheduler job under -the attempt's own deterministic job name; exactly one match repairs the -attempt, recording that job id rather than treating it as job-less, and -an ambiguous match is left alone. Only once an attempt has no -disposition, no scheduler job and no repair, past that age, does -reconcile lock it and record it `lost`, with a null exit code and a -reconcile note explaining why; the unit returns to `ready` while -attempts remain under its allowance, or `failed` otherwise. This is how -a job-less running attempt is resolved rather than left open -indefinitely. - -A retry's own `done_check` -- whether a database-writing stage reuses -an already-complete result set instead of writing a new one -- reads an -attempt's disposition, not only whether its result set is complete. The -rule: a stage's done check reuses a complete, retained result set of -this run with the same kind and provenance key only when that set's -producing attempt is the calling attempt or an attempt whose -disposition is `succeeded`. A set left by an attempt that committed -rows and then failed, or by another attempt still without a -disposition, is not reused: the retry writes a new set under its own -instance, and the orphaned set's rows stay the run's own rows, kept -until the run is deleted, but unreachable because their producer is -never the selected attempt. `load`'s, `crossmatch`'s, `statistics`' and -`prune`'s `done_check` all resolve through this one rule, -`db/sources.find_complete_source_set` and -`db/objects.find_complete_result_set`, each joining `attempts` for the -producing attempt's disposition (the per-stage pages record each -`done_check`'s own key). - -**Instances.** `register` preserves the instance ids and the producing -run, stage and attempt recorded in the manifest, and records the -registering attempt separately. Replaying an identical manifest is a -no-op; conflicting content for an existing instance id is an error. -Database-writing stages create their own attempt-scoped result-set -instances and record completion, including empty completion. After an -uncertain commit the same attempt recovers its recorded results rather -than inserting duplicates. Published identity, content metadata and -provenance are immutable; custody and deletion metadata change only -through promotion and deletion. A dev-registered product enters the run -model by an import run: `reference-import` for a reference `dev` -registered, `psf-import` for a PSF `dev` registered, each a production -run whose one unit registers that product with a run, an attempt and an -instance of its own, so the imported instance is a candidate in project -custody and a production run that depends on it can be promoted: -`dev`'s registered products are the project's, and a scratch import -would refuse every production promotion that depends on it. It can be -named as a frozen input like any other instance. - -**Custody.** Scratch never leaves scratch. Candidate becomes current by -promotion; current becomes candidate again when superseded. Two partial -unique indexes govern the current selection: one holds at most one -current instance per kind and provenance key, for every current row; -the other holds at most one current instance per kind and slot, -wherever the fill has resolved a slot for it ([products](products) -page has the fill). The current selection is a view over the instance -table and includes file products and result sets alike. A consumer -resolves a related selection from one database snapshot. - -**Promotion eligibility.** Automatic and manual promotion require -completed, selected outputs, a recorded image digest identifying a -released artifact, and passing results for every required check under -the resolved check-policy version: an explicit `--check-policy`, then -the run's own `check_policy_ref`, then the default `rebuild-trial@1`; -no unchecked exception exists, so a deliverable of a kind the policy -covers always goes through the gate ([checks](checks) page). Missing or -failed required checks refuse promotion, exit 1. Every provenance -dependency, followed through the whole chain to its roots and not only -the instances named directly, must be `current`, `superseded`, or -itself an after-instance of the same promotion request, validated in -its own right; anything else refuses the promotion, exit 1, naming the -ancestor and its state ([checks](checks) page has the states and the -walk). `run show` prints a `state:` block at the end of its listing, -one line per candidate or current instance in those same states, after -`units:`, `attempts:` and any `promotions:` ([checks](checks) page has -the format). The replacement's kind and slot must equal the requested -selector; it must be retained, and a result set must be complete -([products](products) page has the identity). Every -deliverable's selected producing attempt must have an execution record -whose image digest is a complete release's digest, or the promotion is -refused unless the operator passes the explicit unreleased exception -(the [releases](releases) page has the record and the check). Automatic -promotion stays disabled until the team approves its policy; reprocessing is a production run -with auto-promote off. - -**Promotion.** `promote_run` first fills the run's own rows through -`product_identity_fill()` ([products](products) page), so a candidate a -production run registered before its slot resolved is never refused for -that alone. The default deliverable list is every candidate instance -the run produced through its unit's selected attempt, grouped by (kind, -slot); intermediate revisions and unselected attempts are excluded. A -candidate whose slot, or whose identity, is still null is refused, -naming it, since a slot never exists without its identity ([products](products) -page); two candidates sharing one slot are refused, since a promotion -needs exactly one replacement per kind and slot. Piecemeal -promotion names an explicit subset of that list and passes the same -validation. A change may also name no replacement, withdrawing a slot: -the replaced instance returns to candidate, and the promotion record's -after-instance for that slot is null. -Expected-before per slot is whichever instance currently holds it, or -none; a current instance whose own slot is still null is invisible to -slot promotion, neither replaced nor withdrawn by it, until an operator -resolves it by hand. - -Promotion replaces by slot: a change whose after-instance shares its -predecessor's (kind, slot) supersedes it outright, whatever settings or -upstream instances differ between the two. A reprocessing with changed -settings therefore replaces the old difference images in their slots, -rather than sitting beside them as a new instance. - -A promotion can be planned and frozen ahead of applying it. `run -promote-plan ` reads the run's slot groupings under the lock and -prints them without writing anything; `run promote --plan ` -applies that same file later, under a fresh lock, and refuses, exit 1, -naming the first slot whose actual current instance or whose candidate -no longer matches the plan, writing nothing ([tool](tool) page has the -commands). The plan file itself -must hold a non-empty JSON list of `{kind, slot, before, after}` -entries; anything else, a JSON `null`, an empty list, a different -shape, exits 64 before the file is even read as a plan, let alone -anything written. - -A change into an `association-set` slot whose expected-before is not -null is refused unless the before instance is an ancestor of the after -instance, walked through the provenance key's `base` field, so a live -batch cannot replace a reprocessing campaign's chain head with its own -older chain by accident. Rollback's own recorded inverse is the one -exception: it bypasses this rule, since it is undoing a change that -already passed it once. Ordinary slot -replacement is never a substitute for the chain switch operations.md -describes, the one promotion that moves a whole stream from one catalog -chain to another: that switch is not built, and nothing writes +the attempt's deterministic job name. Exactly one match repairs the +attempt by recording the job id; an ambiguous match is left alone. +Once past that age, an attempt with no disposition, scheduler job or +repair is locked and recorded as `lost`, with a null exit code and a +reconcile note explaining why. Its unit returns to `ready` if attempts +remain under the allowance, or `failed` otherwise. This resolves +job-less running attempts instead of leaving them open indefinitely. + +### Reusing result sets + +A database-writing stage's `done_check` decides whether to reuse a +complete result set or write a new one. It checks the producing +attempt's disposition as well as completeness. Reuse requires a +complete, retained set in this run with the same kind and provenance +key, produced either by the calling attempt or by an attempt whose +disposition is `succeeded`. + +A set left by an attempt that committed rows and then failed, or by +another attempt still without a disposition, is not reused. The retry +writes a new set under its own instance. The orphaned rows remain the +run's own until deletion, but are unreachable because their producer +is never the selected attempt. + +`load`'s, `crossmatch`'s, `statistics`' and `prune`'s `done_check` +all use this rule through `db/sources.find_complete_source_set` and +`db/objects.find_complete_result_set`. Both join `attempts` for the +producing attempt's disposition; the per-stage pages record each +`done_check`'s key. + +## Instances and custody + +`register` preserves the instance ids and producing run, stage and +attempt from the manifest, recording the registering attempt +separately. Replaying an identical manifest is a no-op; conflicting +content for an existing instance id is an error. Database-writing +stages create their own attempt-scoped result-set instances and record +completion, including empty completion. After an uncertain commit, the +same attempt recovers its recorded results instead of inserting +duplicates. + +Published identity, content metadata and provenance are immutable. +Custody and deletion metadata change only through promotion and +deletion. Scratch never leaves scratch. Promotion makes a candidate +current; supersession returns it to candidate. + +A product registered by `dev` enters the run model through an import +run: `reference-import` for a reference, `psf-import` for a PSF. Each +is a production run whose one unit registers the product with its own +run, attempt and instance. The instance is a candidate in project +custody and can be named as a frozen input. Production runs that +depend on it can therefore be promoted: `dev`'s registered products +belong to the project, while a scratch import would refuse every +dependent production promotion. + +Two partial unique indexes govern current selection. One allows at +most one current instance per kind and provenance key across all +current rows. The other allows at most one per kind and slot wherever +the fill has resolved a slot ([products](products) page). Current +selection is a view over the instance table, covering file products +and result sets alike. A consumer resolves a related selection from +one database snapshot. + +## Promotion + +### Eligibility + +Automatic and manual promotion require completed, selected outputs, +a recorded image digest identifying a released artifact, and passing +results for every required check under the resolved check-policy +version. Policy resolution uses an explicit `--check-policy`, then the +run's `check_policy_ref`, then the default `rebuild-trial@1`. There is +no unchecked exception: every deliverable of a kind the policy covers +passes the gate ([checks](checks) page). Missing or failed required +checks refuse promotion with exit 1. + +Every provenance dependency, followed to its roots through the whole +chain, must be `current`, `superseded`, or an after-instance of the +same promotion request that passes validation in its own right. +Anything else refuses promotion with exit 1, naming the ancestor and +its state ([checks](checks) page has the states and walk). `run show` +ends its listing with a `state:` block, one line per candidate or +current instance in those states, after `units:`, `attempts:` and any +`promotions:` ([checks](checks) page has the format). + +The replacement's kind and slot must match the requested selector. +It must be retained, and a result set must be complete +([products](products) page has the identity). Every deliverable's +selected producing attempt must have an execution record whose image +digest belongs to a complete release. Otherwise promotion is refused +unless the operator passes the explicit unreleased exception +(the [releases](releases) page has the record and check). Automatic +promotion stays disabled until the team approves its policy. +Reprocessing is a production run with auto-promote off. + +### Selecting replacements + +`promote_run` first fills the run's rows through +`product_identity_fill()` ([products](products) page). A candidate +registered before its slot resolved is therefore not refused for that +alone. The default deliverable list contains every candidate the run +produced through its unit's selected attempt, grouped by (kind, slot). +It excludes intermediate revisions and unselected attempts. + +A candidate whose slot or identity remains null is refused by name: +a slot never exists without its identity ([products](products) page). +Two candidates sharing a slot are refused because each (kind, slot) +needs exactly one replacement. Piecemeal promotion names an explicit +subset of the list and passes the same validation. + +A change may withdraw a slot by naming no replacement. The replaced +instance returns to candidate, and the promotion record's after-instance +is null. Expected-before for a slot is its current instance, or none. +A current instance with a null slot is invisible to slot promotion: +it can be neither replaced nor withdrawn until an operator resolves it +by hand. + +A replacement supersedes its predecessor whenever they share (kind, +slot), regardless of differing settings or upstream instances. +Reprocessing with changed settings therefore replaces old difference +images in their slots instead of adding instances beside them. + +### Frozen plans + +`run promote-plan ` reads the run's slot groupings under the lock +and prints them without writing anything. Later, +`run promote --plan ` applies that file under a fresh lock. +It refuses with exit 1, naming the first slot whose actual current +instance or candidate no longer matches the plan, and writes nothing +([tool](tool) page has the commands). + +The plan file must contain a non-empty JSON list of +`{kind, slot, before, after}` entries. A JSON `null`, an empty list or +any other shape exits 64 before the file is read as a plan or anything +is written. + +### Applying changes + +A change into an `association-set` slot with a non-null expected-before +is refused unless that instance is an ancestor of the after-instance, +following the provenance key's `base` field. This prevents a live +batch from replacing a reprocessing campaign's chain head with its +own older chain. Rollback's recorded inverse is the sole exception: +it undoes a change that already passed this rule. + +Ordinary slot replacement cannot substitute for the chain switch +operations.md describes, the promotion that moves a whole stream +between catalog chains. That switch is not built, and nothing writes `loop_dates.kind = 'switch'`. -Only a production run's outputs can be promoted; scratch never leaves -scratch. On rows a run wrote, `vbest` is a current-membership flag: 1 -while the instance is current, 0 otherwise. Rows `dev` wrote, with `run` -null, including those an import run links through `instance`, keep -`dev`'s own flag. A mapped kind whose instance has no row refuses the -promotion. All promotions take one transaction-scoped advisory lock; -after acquiring it the transaction checks every expected previous -selection, including expected absence, against the actual selection and -refuses the whole request on any mismatch, then validates dependencies, -updates custody and records the action. Each promotion records a before -and after instance for every affected slot, either nullable, and -`promotion_changes` carries the slot alongside the provenance key. - -Rollback inverts the exact recorded change, by slot: the recorded +Only production outputs can be promoted; scratch never leaves +scratch. On rows a run wrote, `vbest` is 1 while the instance is +current and 0 otherwise. Rows `dev` wrote with `run` null keep +`dev`'s flag, including rows an import run links through `instance`. +A mapped kind whose instance has no row refuses promotion. + +Every promotion takes one transaction-scoped advisory lock. Under it, +the transaction checks all expected previous selections, including +expected absence, against the actual selection. Any mismatch refuses +the whole request. It then validates dependencies, updates custody and +records the action. Each affected slot has a before and after instance, +either nullable; `promotion_changes` stores the slot alongside the +provenance key. + +### Rollback + +Rollback inverts the exact recorded change by slot: the recorded after-instance becomes the expected before, and the recorded before-instance, possibly null, becomes the new after. A change recorded before slot identity, with no slot, is refused as not reversible ({ref}`legacy compatibility ruling `); every promotion `rapid_rebuild` has -made was recorded after slot identity, so the refusal is reachable only -against a disposable database. Reversal is itself a promotion, carrying -the inverse mapping and `request_context.rollback_of`, and is refused if -the recorded after-selection is no longer current. `promote()` accepts -only a slot selector; anywhere else a selector naming an instance whose -own slot is null is refused. - -**Deletion.** `run delete` is allowed on a scratch run by its owner, -finished or not: a finished run admits no new unit, attempt or input -binding, but that alone does not block its deletion. It first locks the run, verifies the owner, refuses if any -attempt is queued, running or unresolved, meaning it carries no -disposition at all; `lost` is itself a recorded resolution of that -uncertainty, written only once reconcile finds the scheduler no longer -returns the job, or once `run reconcile --resolve-jobless` finds no job -under the attempt's name, so a `lost` attempt does not by itself block -deletion, if any frozen input binding -of unfinished work or any provenance dependency of a retained output -outside the run points into it, or if a row in `xsources` references -one of the run's rows (that table is not in the cleanup set below, so -such a reference refuses the whole delete rather than leaving an orphan), -or if a row in `refimimages`, -`refimcatalogs` or `refimmeta` belonging to a *different* run's -reference references one of the run's rows (its own reference's -satellite rows are removed by cleanup below, not refused), and marks the run deleting, in one transaction. Attempt -allocation, input binding and result acceptance use the same run fence -and refuse a deleting or deleted run. - -The guard's count of blocking consumers is live consumers only: it -counts a `unit_inputs` binding from outside the run only when the -binding unit's own run is in a state other than `deleted`, and a -`dependencies` edge from outside the run only when the consumer -instance's `deletion_state` is other than `deleted`. Tombstone rows are -never removed; a consumer run that is still `deleting` still blocks, -and only a consumer run that has finished deleting stops counting -against the run that produced what it once consumed. The refusal itself -is exit 64 from `run delete`. This is what makes the input binding above safe to write -at submission rather than at output registration: a unit's frozen -bindings, recorded through `bind_unit_inputs` before its attempt starts, -make it a live consumer of its declared inputs from that point, and the -guard sees that consumer as soon as it exists, not only once it has -produced something of its own to depend on. - -Before touching storage, every attempt's output location must lie in +made was recorded after slot identity, so this refusal is reachable +only against a disposable database. + +Reversal is itself a promotion, carrying the inverse mapping and +`request_context.rollback_of`. It is refused if the recorded +after-selection is no longer current. `promote()` accepts only a slot +selector; anywhere else, a selector naming an instance whose own slot +is null is refused. + +## Deletion + +### Guarding the run + +An owner may `run delete` a scratch run whether or not it is finished. +A finished run admits no new unit, attempt or input binding, but can +still be deleted. In one transaction, deletion locks the run, verifies +the owner, checks for blockers and marks the run deleting. + +Deletion is refused if: + +- An attempt is queued, running or unresolved, meaning it has no + disposition. +- A frozen input binding of unfinished work or a provenance dependency + of a retained output outside the run points into it. +- A row in `xsources` references one of the run's rows. That table is + outside the cleanup set, so deletion would leave an orphan. +- A row in `refimimages`, `refimcatalogs` or `refimmeta` belonging + to another run's reference points to one of this run's rows. Its own + reference's satellite rows are removed during cleanup. + +`lost` is a recorded resolution of uncertainty. Reconcile writes it +only after the scheduler no longer returns the job, or after +`run reconcile --resolve-jobless` finds no job under the attempt's +name. A `lost` attempt therefore does not by itself block deletion. +Attempt allocation, input binding and result acceptance use the same +run fence and refuse a deleting or deleted run. + +The guard counts only live consumers outside the run. A `unit_inputs` +binding counts only when its unit's run is not `deleted`; a +`dependencies` edge counts only when the consumer instance's +`deletion_state` is not `deleted`. Tombstones are never removed. +A consumer run still `deleting` blocks deletion of its producer's run +until the consumer finishes deleting. `run delete` refuses with exit 64. + +Bindings recorded through `bind_unit_inputs` before an attempt starts +make the unit a live consumer immediately. The deletion guard sees it +before it produces any output, which makes submission-time binding +safe. + +### Storage and row cleanup + +Before storage is touched, every attempt's output location must lie in the scratch bucket under `runs//`; otherwise the whole delete -is refused. Cleanup then removes the run's object versions; a storage -delete that reports a per-object error leaves the run deleting and -touches no row, so a retry resumes at storage. Once storage cleanup is -clean, one transaction removes the science rows, marks the run's -instance rows deleted (`deletion_state`) and marks the run deleted. -The science-row cleanup set is -`l2files`, `l2filemeta`, `refimages`, `diffimages`, `diffimmeta`, -`psfs` and `sources` (the parent delete reaches the children), each -where the `run` column equals the run being deleted; `l2files` and -`l2filemeta` join the set because a scratch run's admitted rows are -its own, and the fence already refuses when another run binds them -`refimimages`, `refimcatalogs` and -`refimmeta` carry no `run` column of their own (they are reached only -through the `refimages` row's `rfid`), so the same transaction deletes -them first, joined through this run's own `refimages` rows, before -deleting those `refimages` rows; a reference belonging to another run is -untouched, and the refusal check above already keeps this run's rows -from being deleted out from under it. Rows with `run` NULL, which is -everything `dev` wrote, are never touched. Failure -before the final transaction leaves the run deleting, and re-running -cleanup on a deleting run resumes it from wherever it stopped. Custody stays separate from deletion -state. Run, unit, attempt and instance rows remain as tombstones with -their provenance. Scratch expiry calls the same operation after the -run's expiry date, with a warning first and a pin to hold a run. +is refused. Cleanup removes the run's object versions first. A +per-object storage deletion error leaves the run deleting and touches +no row; a retry resumes at storage. + +After clean storage removal, one transaction removes the science rows, +marks the instance rows deleted (`deletion_state`) and marks the run +deleted. The science-row cleanup set is `l2files`, `l2filemeta`, +`refimages`, `diffimages`, `diffimmeta`, `psfs` and `sources` +(the parent delete reaches the children), each where `run` equals the +run being deleted. `l2files` and `l2filemeta` are included because a +scratch run owns its admitted rows, and the fence already refuses +deletion when another run binds them. + +`refimimages`, `refimcatalogs` and `refimmeta` have no `run` column: +they are reached through the `refimages` row's `rfid`. The same +transaction joins them through this run's `refimages` rows and deletes +them before those parent rows. Another run's reference is untouched; +the guard prevents deleting rows it references. Rows with `run` NULL, +everything `dev` wrote, are never touched. + +Failure before the final transaction leaves the run deleting. +Re-running cleanup resumes where it stopped. Custody stays separate +from deletion state. Run, unit, attempt and instance rows remain as +tombstones with their provenance. Scratch expiry calls the same +operation after the expiry date, with a warning first and a pin to +hold a run. ## Storage layout -Two buckets separate personal and project custody: `s3:///rapidpipe` -for scratch runs, `s3:///rapidpipe` for production -runs. The launcher chooses the root from the run's kind at submission. -A scratch run reads `RAPIDPIPE_OUTPUTS_ROOT_SCRATCH`, falling back to -`RAPIDPIPE_OUTPUTS_ROOT` when it is unset. A production run requires -`RAPIDPIPE_OUTPUTS_ROOT_PRODUCTION`; it never falls back to the -unsuffixed variable, which is the scratch fallback only. Runs written -under `rapidpipe-firstrun/` stay where they were written; their attempt -rows carry that location rather than moving to the new root. Candidate and current objects share the project bucket and -never move at promotion. Ordinary execution roles cannot delete -completed objects; only the cleanup role deletes, and only through -`run delete`. Within either bucket: +Two buckets separate personal and project custody: +`s3:///rapidpipe` for scratch runs and +`s3:///rapidpipe` for production runs. At submission, +the launcher chooses the root from the run's kind. Scratch reads +`RAPIDPIPE_OUTPUTS_ROOT_SCRATCH`, falling back to +`RAPIDPIPE_OUTPUTS_ROOT` when unset. Production requires +`RAPIDPIPE_OUTPUTS_ROOT_PRODUCTION` and never falls back to the +unsuffixed variable. + +Runs written under `rapidpipe-firstrun/` stay there; their attempt +rows retain that location. Candidate and current objects share the +project bucket and never move at promotion. Ordinary execution roles +cannot delete completed objects. Only the cleanup role deletes, and +only through `run delete`. Within either bucket: ``` runs/////manifest.json runs///// ``` -Admitted inputs and reference images live under the same scheme in the -project bucket, in the run that admitted or built them. Consumers read -current products through the selection, never by guessing paths. Alert -containers follow the same scheme; the outbox row carries the -container's location and each alert's block locator, the offset and -size of the Avro block holding its record plus the record's position -within that block and within the container, since Avro compresses -records per block and a single record has no byte range of its own -(the [alerts](alerts) page has the outbox's full column list). - -Each kind runs under its own Batch job definition. Scratch runs use -`rapid-rebuild`, whose job role can write only the scratch bucket; -production runs use `rapid-rebuild-production`, whose job role can -write `rapidpipe/*` of the products bucket and never delete, and also -reads the scratch bucket's `rapidpipe*` prefixes, where staged inputs, -settings overlays and input-set manifests live. A scratch run reads `RAPIDPIPE_BATCH_JOB_DEFINITION_SCRATCH`, -falling back to `RAPIDPIPE_BATCH_JOB_DEFINITION` when it is unset. A -production run requires `RAPIDPIPE_BATCH_JOB_DEFINITION_PRODUCTION`; it -never falls back to the unsuffixed variable. - -Deletion runs under different credentials than execution. `run delete` -runs launcher-side, under the caller's own credentials, not a Batch -job's. On a workstation, the instance role holds `rapid-scratch-cleanup`, -which can delete object versions under the scratch bucket's -`rapidpipe*/runs/*` prefix and nothing else; the deletion code itself -refuses to touch any object outside the scratch bucket. The dedicated -cleanup principal for workstation submission is on the [tool](tool) -page. +Admitted inputs and reference images use this scheme in the project +bucket, in the run that admitted or built them. Consumers read current +products through the selection, never by guessing paths. Alert +containers use the same scheme. The outbox row holds the container's +location and each alert's block locator: the offset and size of its +Avro block, plus the record's position within the block and container. +Avro compresses records per block, so a single record has no byte range +of its own (the [alerts](alerts) page has the outbox's full column list). + +### Execution and cleanup credentials + +Each run kind has its own Batch job definition. Scratch uses +`rapid-rebuild`, whose job role writes only to the scratch bucket. +Production uses `rapid-rebuild-production`, whose job role can write +`rapidpipe/*` in the products bucket but never delete. It also reads +the scratch bucket's `rapidpipe*` prefixes, which hold staged inputs, +settings overlays and input-set manifests. + +Scratch reads `RAPIDPIPE_BATCH_JOB_DEFINITION_SCRATCH`, falling back to +`RAPIDPIPE_BATCH_JOB_DEFINITION` when unset. Production requires +`RAPIDPIPE_BATCH_JOB_DEFINITION_PRODUCTION` and never falls back to the +unsuffixed variable. + +`run delete` runs launcher-side under the caller's credentials, not a +Batch job's. On a workstation, the instance role holds +`rapid-scratch-cleanup`, which can delete object versions only under +the scratch bucket's `rapidpipe*/runs/*` prefix. The deletion code +refuses any object outside the scratch bucket. The [tool](tool) page +names the dedicated cleanup principal for workstation submission. ## Identifiers Runs, attempts and product instances receive globally unique, time-ordered identifiers (ULID) before execution or manifest -publication, allocated by whoever creates the row, so the local runner -and the launcher need no database to allocate; registration -preserves them. Date-and-sequence names such as `r-20260921-007` are -display labels only. Unit identity is unique within a run and stage and -uses the product vocabulary's unit identifiers. +publication. Whoever creates the row allocates the id; the local +runner and launcher need no database for allocation, and registration +preserves it. Date-and-sequence names such as `r-20260921-007` are +display labels only. Unit identity is unique within a run and stage +and uses the product vocabulary's unit identifiers. ## Schema -The run-model tables above land as one migration beside the baseline. -The additive columns and widened keys on the `dev` product tables land -one kind at a time, difference image first, each with the stage that -writes it. The `dev` schema coexists and is kept as far as possible; the rebuild adds, it does not rename or drop. +The run-model tables land as one migration beside the baseline. +Additive columns and widened keys on the `dev` product tables land +one kind at a time, difference image first, with the stage that writes +each kind. The `dev` schema coexists and is kept as far as possible; +the rebuild adds without renaming or dropping. ## Trial database A scratch run's `database target` may name a trial database instead of -the one production database. A trial database lives on the production -PostgreSQL instance, beside `rapid`, created with rapid's own schema: -it is not a separate server or a clone. rapid's own migration applier -(`database/apply-migrations.sh`, `database/migrations/`, files named -`YYYYMMDD-NN-name.sql`) builds and updates it, run inside the pipeline -container image; the applier records each applied file's sha256 and an -applied migration is never edited. Clients reach a trial database -through the pooler: pgbouncer routes any database name to the one -PostgreSQL port (a wildcard `[databases]` entry), so adding a trial -database needs no pooler config change. The pipeline authenticates as a -per-tier service login (for example `rapid_rebuild_pipeline`) holding -SELECT/INSERT/UPDATE/DELETE on the trial database's tables, granted by -guarded migrations in rapid's own stream (they no-op where the role -does not exist, so CI still applies the stream from empty); people read -it through `rapid_read`. That role's identity (its NOLOGIN cluster -role, its secret and its SSM tree) is provisioned separately, by the -system repo's own migration stream, since the role itself is -cluster-wide and rapid's stream owns only the trial database's tables. -A job definition's `RAPID_PARAMETER_PATH` selects the tier: it points -at an SSM tree `/rapid//db/*` (name, host, port, secret id) that -resolves the trial database's name, host and port, with the service -login's password in Secrets Manager under -`rapid/db/service/-pipeline`. Each tier runs under its own Batch -job definition, revision-pinned to an image digest. The rebuild's own -tier is trial database `rapid_rebuild`, service login -`rapid_rebuild_pipeline`, tree `/rapid/rebuild`, job definition -`rapid-rebuild`. A run created under a release submits to the revision -that release's record deployed for the run's kind, not whatever -revision the job definition currently pins (the [releases](releases) -page has the mechanism). +the one production database. A trial database lives beside `rapid` +on the production PostgreSQL instance, built with rapid's own schema. +It is neither a separate server nor a clone. + +rapid's migration applier (`database/apply-migrations.sh`, +`database/migrations/`, files named `YYYYMMDD-NN-name.sql`) builds +and updates it from inside the pipeline container image. The applier +records each file's sha256; an applied migration is never edited. +Clients connect through pgbouncer, whose wildcard `[databases]` entry +routes any database name to the one PostgreSQL port. Adding a trial +database therefore needs no pooler config change. + +The pipeline authenticates as a per-tier service login, for example +`rapid_rebuild_pipeline`, with SELECT/INSERT/UPDATE/DELETE on the trial +database's tables. Guarded migrations in rapid's stream grant these +permissions; they no-op where the role does not exist, so CI can apply +the stream from empty. People read through `rapid_read`. That role's +identity (its NOLOGIN cluster role, secret and SSM tree) is provisioned +separately by the system repo's migration stream: the role is +cluster-wide, while rapid's stream owns only the trial database's tables. + +A job definition's `RAPID_PARAMETER_PATH` selects the tier. It points +at the SSM tree `/rapid//db/*` (name, host, port, secret id), +which resolves the trial database's name, host and port. The service +login's password is in Secrets Manager under +`rapid/db/service/-pipeline`. Each tier has its own Batch job +definition, revision-pinned to an image digest. The rebuild tier uses +database `rapid_rebuild`, service login `rapid_rebuild_pipeline`, +tree `/rapid/rebuild` and job definition `rapid-rebuild`. + +A run created under a release submits to the revision that release's +record deployed for the run's kind, regardless of what revision the +job definition currently pins (the [releases](releases) page has the +mechanism). ## Python interface