Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
280 changes: 128 additions & 152 deletions system/checks.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,55 +2,38 @@

**Status: DRAFT**

What a check is, the two checks the rebuild ships, how a check policy is
defined and versioned, how promotion validates against one, and
automatic promotion's design. The [runs](runs) page's Promotion
eligibility paragraph and the [specification](specification)'s
Promotion section state the requirement; this page records how the
rebuild meets it.

## In plain terms

A check is a function that looks at one product instance already in
the database and records whether it passes. A check policy is a named,
versioned list of checks with their bounds, some required and some
advisory, and it carries its own approval level: unapproved policies
gate nothing, a trial-approved policy can gate a person's promotion,
and only a team-approved one can gate an automatic one. Promotion runs
the policy's required checks and refuses if any is missing or failed;
nothing about this page changes what promotion already does with a
passing or failing check, it defines what running one means. Automatic
promotion is built and tested, and stays off in production until the
team approves a policy for it.

Plainly: `rebuild-trial@1`, the policy every promotion in this rebuild
has run under so far, carries a trial approval, not the team's own
sign-off. That trial approval is enough to gate a person's `run promote`
by hand, which is what every demonstration on this page and the
[runs](runs) page has done. It is not enough to gate an automatic
promotion, and no policy the rebuild ships is. The team's own approval
of a policy, and whether any policy the team approves also permits
automatic promotion, are both still open.
A check examines one product instance in the database and records
whether it passes. A check policy names and versions a list of required
and advisory checks with their bounds. Promotion runs the required
checks and refuses if any is missing or failed. This page defines how
checks run without changing how promotion treats their results, meeting
the requirement in the [runs](runs) page's Promotion eligibility
paragraph and the [specification](specification)'s Promotion section.

Every promotion in the rebuild so far, including every demonstration on
this page and the [runs](runs) page, has used `rebuild-trial@1` for a
person's `run promote` by hand. Its trial approval permits manual
promotion. Automatic promotion is built and tested but stays off in
production: no shipped policy has the team's approval to permit it.

## What a check is

A check is a named, versioned Python function over one product
instance, registered in `rapidpipe/checks/`: a registry keyed
`name@version`, declared with `@check("difference-image-statistics",
"1")`. Its signature is `(conn, instance_id, params: dict) ->
CheckResult(outcome 'passed'|'failed', detail dict)`. Running a check
always records a `checks` row: instance, check name, version, the
required flag from the policy that ran it, outcome, and a detail
document of the measurements, the bounds, the params it ran under and
the reason for the outcome; a check that raises records `failed` with
the error in `detail`, never nothing. Recording the params matters
because a policy can override a shipped check's bounds: a result
recorded under other params is not a result under these, so the
promotion gate below only ever counts a row whose recorded params equal
the resolving policy's params for that check. Versions are strings, and
a changed threshold or a changed measurement is a new version, not an
edit to an existing one. A check reads the database; it may read S3
later, but both checks this page ships are database-only.
instance. It is registered in `rapidpipe/checks/`, keyed by
`name@version`, with a declaration such as
`@check("difference-image-statistics", "1")`. Its signature is
`(conn, instance_id, params: dict) -> CheckResult(outcome 'passed'|'failed', detail dict)`.
Versions are strings; a changed threshold or measurement requires a
new version. Both shipped checks read only the database; future checks
may read S3.

Running a check always records a `checks` row: instance, check name,
version, the policy's required flag, outcome, and a detail document
containing the measurements, bounds, params and reason for the outcome.
A check that raises records `failed` with the error in `detail`.
Policies can override a shipped check's bounds, so the promotion gate
counts only rows whose recorded params match the resolving policy's
params for that check.

## The two checks

Expand All @@ -65,24 +48,20 @@ ratio falls within `[ratio_lo, ratio_hi]`. All bounds come from the
policy's `params` for that check.

`catalog-counts-vs-reference@1` runs over a `source-set` result-set
instance. It now finds its reference by slot rather than walking
dependencies by hand: the candidate's slot already names its exposure,
detector and catalog type ([products](products) page), so the check
looks for another current `source-set` instance sharing that same slot,
any differencer, or, when `params` names a `reference_run`, that run's
instance for the same slot instead. A candidate or a would-be reference
without a slot is not a reference; the check's own runner fills every
candidate's slot first, so a set only misses one when its producer is
still unresolved (below). It compares the candidate's
instance. It finds its reference by slot instead of walking dependencies
by hand. The slot names the candidate's exposure, detector and catalog
type ([products](products) page). The reference is another current
`source-set` instance in the same slot, from any differencer, or, when
`params` names a `reference_run`, that run's instance for the same slot.
A candidate or would-be reference without a slot is not a reference.
The runner fills every candidate's slot first, so a set lacks one only
when its producer is still unresolved. The check compares the candidate's
`result_sets.row_count` against the reference's and passes when
`|candidate − reference| / reference` is at most `params.tolerance`.
When no reference instance exists, the outcome follows
`params.missing_reference`, recorded in `detail`; policy
`rebuild-trial@1` sets that to pass.

Policy `rebuild-trial@1` marks `difference-image-statistics` required
and `catalog-counts-vs-reference` advisory, required false.

## Check policies

A check policy is a named, versioned TOML file shipped in the package,
Expand All @@ -91,133 +70,130 @@ same `name@version` key as a check. Its fields are `name`, `version`,
`approval` (`none`, `trial` or `team`), `approved_by` (nullable; who
granted the recorded approval level), `auto_promote` (boolean), and a
`[[checks]]` table per check listing its name, version, kind, required
flag and params. A policy is immutable once landed; a change to any of
these is a new version, never an edit in place. Promotion, manual or
automatic, refuses a policy whose `approval` is `none`; automatic
promotion additionally requires `team` and `auto_promote` true, so a
`trial`-approved policy can gate a person's promotion but never an
automatic one.

One policy ships with the rebuild. `rebuild-trial@1`'s bounds are set
from the control run's measured values with generous margins:
flag and params. A policy is immutable once landed; changing any field
requires a new version.

Manual and automatic promotion refuse a policy whose `approval` is
`none`. A `trial`-approved policy can gate a person's promotion;
automatic promotion requires `team` and `auto_promote` true.

The rebuild ships one policy, `rebuild-trial@1`. It marks
`difference-image-statistics` required and `catalog-counts-vs-reference`
advisory, required false. Its bounds come from the control run's measured
values with generous margins:
`scalefacref` in `[1e-3, 1e5]`, `dxrmsfin` and `dyrmsfin` each at most
2.0 pixels, `|dxmedianfin|` and `|dymedianfin|` each at most 1.0 pixel,
`nsexcatsources` in `[1000, 1e6]`, and the SExtractor positive-to-negative
ratio in `[0.1, 10]`. `rebuild-strict@1`, the same two checks with
bounds the control run cannot meet (`scalefacref` in `[0.99, 1.01]`,
the same two RMS measurements each at most 0.01 pixel, `nsexcatsources`
at most 1000), is a test fixture that exercises a refusal; it does not
ship, and naming it to a command is an unknown policy. The shipped
policy carries `approval: trial`, `approved_by` recording who
granted it, and `auto_promote` false. This trial approval
satisfies the mechanism the runs and specification pages ask for well
enough to demonstrate manual promotion by hand; it is not the team's own
sign-off, which is a recorded open item, and it does not permit
automatic promotion regardless. There is no policy table in the
database: the promotion row records the policy's version string and the
ids of the check rows it relied on, which is the durable record of what
ran; the TOML file is the versioned definition of what the policy meant
at that version.
ratio in `[0.1, 10]`. It carries `approval: trial`, `approved_by`
recording who granted it, and `auto_promote` false. This trial approval
suffices to demonstrate the manual-promotion mechanism required by the
runs and specification pages; the team's own sign-off remains open.

`rebuild-strict@1` is a refusal test fixture using the same two checks
with bounds the control run cannot meet: `scalefacref` in
`[0.99, 1.01]`, the same two RMS measurements each at most 0.01 pixel,
and `nsexcatsources` at most 1000. It does not ship; naming it to a
command is an unknown policy.

There is no policy table in the database. The promotion row durably
records what ran through the policy's version string and the ids of the
check rows it relied on. The TOML file defines what that version meant.

## The check commands

`run_policy_checks`, the runner called by both `check run` and the
promotion gate, fills every candidate's slot before running any check.
Thus `catalog-counts-vs-reference` never sees a slot the fill could have
resolved. Checks run on the workstation and reach the database through
the instance role; they do not touch the pipeline image. The [tool](tool)
page gives the full command surface for `check list`, `check run` and
`check show`, including every flag, print format and exit code.

## The promotion gate

`promote()` takes an optional check policy. `promote_run` and `run
promote` resolve which one to validate against in this order: an
explicit `--check-policy`, then the run's own `check_policy_ref`, then
the default `rebuild-trial@1`. Promotion first refuses an unapproved
policy (`approval: none`) before looking at any check result. Every
promotion under an approved policy is then validated under its version:
for each after-instance, every policy check whose kind matches that
instance must have a latest `checks` row, for that check's name,
version and the policy's own params for it, with outcome `passed` when
the policy marks it required; ties in `happened_at` break by row id, and
the read locks the rows it relies on, so a check recorded after the read
started is not counted. A missing or failed required check refuses the
whole promotion, exit 1, naming the instance, the check and the
outcome in the message; a failed or missing advisory check is recorded
but does not refuse. The promotion row records `check_policy_version`
and `check_result_ids`, the rows it relied on. A kind the policy names
no check for passes trivially, and there is no unchecked-exception
flag: a deliverable of a kind the policy does cover always goes through
this gate.

Promotion also walks every dependency to its roots, not only the
after-instances themselves. `_refuse_unpromoted_ancestor` follows
`dependencies` recursively for each after-instance, cycle-guarded
rather than depth-capped, reading the same instance states `run show`
prints ([runs](runs) page, "Rules"). An ancestor that is itself
`current` or `superseded` passes, since its own chain was judged when
it was promoted, but the walk continues past it rather than stopping
there, so a `candidate` grandparent behind a current parent still
refuses the promotion. An ancestor that is itself an after-instance of
the same promotion request also passes the walk, but gets no exemption
from validation: it is judged in its own right, against the same
eligibility rule and the same check policy every after-instance
answers to. Any ancestor that is neither `current` nor `superseded` nor
a member of the request's own after-instances refuses the whole
promotion, exit 1, naming the after-instance, the ancestor, its kind
and its state: "after instance '01...' depends on '01...'
the default `rebuild-trial@1`. Promotion refuses an unapproved policy
(`approval: none`) before reading any check result.

For each after-instance, every check in the approved policy whose kind
matches the instance must have a latest `checks` row for its name,
version and policy params. Required checks must have outcome `passed`.
Ties in `happened_at` break by row id. The read locks the rows it relies
on; a check recorded after the read starts is not counted.

A missing or failed required check refuses the whole promotion, exit 1,
with a message naming the instance, check and outcome. A failed or
missing advisory check is recorded but does not refuse. The promotion
row records `check_policy_version` and `check_result_ids`, the rows it
relied on. A kind with no check in the policy passes trivially. There is
no unchecked-exception flag: a deliverable of a kind covered by the
policy always goes through this gate.

Acceptance means required checks pass, nothing more ({ref}`acceptance
ruling <decision-acceptance>`). There is no separate acceptance record:
the gate re-reads the latest result of every required check each time.
The `checks` row is the durable record of what let a candidate through;
no judgement is stored on the instance.

### Dependencies and rollback

Promotion walks every dependency to its roots.
`_refuse_unpromoted_ancestor` follows `dependencies` recursively for
each after-instance, with a cycle guard and no depth cap. It reads the
same instance states that `run show` prints ([runs](runs) page,
"Rules"). A `current` or `superseded` ancestor passes because its chain
was judged when it was promoted. The walk still continues past it, so a
`candidate` grandparent behind a current parent refuses promotion.

An ancestor included among the same request's after-instances also
passes the walk, but must independently meet the same eligibility rule
and check policy as every after-instance. Any ancestor that is neither
`current` nor `superseded` nor an after-instance in the request refuses
the whole promotion, exit 1. The message names the after-instance,
ancestor, kind and state: "after instance '01...' depends on '01...'
(difference-image, candidate: not current or superseded; promote it
first or in the same request); refusing", the hint appearing only for
a `candidate` ancestor. A refusal changes no selection or custody; the
transaction raises before anything is written.

Rollback skips both gates, the check-policy gate and the ancestor walk:
it restores a selection an earlier promotion already admitted, the same
treatment the released-image check gives it.
Each direct dependency of the restored selection must still be complete,
retained and in project custody; only the walk past those direct
dependencies is skipped.
Rollback skips the check-policy gate and the ancestor walk because it
restores a selection an earlier promotion admitted, the same treatment
the released-image check gives it. Each direct dependency of the
restored selection must still be complete, retained and in project
custody; only the walk past those direct dependencies is skipped.

### Accepted defect

The gate's read of the latest check rows takes its lock `FOR SHARE`; a
failed check row committed after that read is not seen by the
promotion it should have refused. The dependency walk reads no check
rows: an ancestor's state is custody, completeness and selection. This
is an accepted defect, not fixed.

## Acceptance

Acceptance is required checks pass, nothing more ({ref}`acceptance
ruling <decision-acceptance>`). There is no acceptance record separate
from a `checks` row: the promotion gate above re-reads the latest
result of every required check each time it runs, and that check
result, not a stored judgement on the instance, is the durable record
of what let a candidate through.

## Automatic promotion

Automatic promotion is designed in and switched off. `run create
--auto-promote [--check-policy P]` is accepted only when the named
`run create --auto-promote [--check-policy P]` is accepted only when the named
policy carries the team's approval, recorded as `approval: team` in the
policy file, and its `auto_promote` is true; otherwise it is refused,
exit 1, with a message naming the policy and stating that team
approval is pending. The shipped policy does not carry the
team's approval, so it does not permit it. At the end
of a `run start` walk, once every unit is complete, the tool calls
`maybe_auto_promote(conn, run_id)`: with the run's `auto_promote` flag
approval is pending.

At the end of a `run start` walk, once every unit is complete, the tool
calls `maybe_auto_promote(conn, run_id)`: with the run's `auto_promote` flag
true, it runs the resolved policy's checks over the run's candidates and
promotes on a pass; otherwise it prints `auto-promote off (policy
<ref>)` and does nothing further, immediately before `start`'s own
final `run=<id> state=complete` line. Every run created so far prints
this line, since no shipped policy permits automatic promotion. The
path is exercised end to end against a fixture policy in tests; in
production it can currently only print, because no policy carries the
approval that would let it act.

## The check commands

`run_policy_checks`, the runner both `check run` and the promotion gate
call, fills every candidate's slot before running a single check, so
`catalog-counts-vs-reference` never sees a slot the fill itself could
have resolved. Checks run on the workstation, reaching the database
through the instance role; none of this touches the pipeline image. The
full command surface, `check list`, `check run` and `check show`,
including every flag, print format and exit code, is on the
[tool](tool) page.
this line, since no shipped policy permits automatic promotion.
Tests exercise the path end to end with a fixture policy. In production
it can currently only print: no policy carries the approval to act.

## Not decided here

- The team's own sign-off on a policy, beyond `rebuild-trial@1`'s trial
approval, and the content of any policy the team approves for
automatic promotion: both scientific, and the team's.
approval, whether an approved policy permits automatic promotion, and
the content of any policy approved for it. These are scientific
questions for the team.
- Checks that read S3 rather than only the database.
Loading