From 780e704f3ca0f3fc0b31f894562117ef547b697c Mon Sep 17 00:00:00 2001 From: Ben Rusholme Date: Mon, 28 Sep 2026 11:37:12 -0700 Subject: [PATCH] statistics, checks: clarity pass (Codex) Co-Authored-By: Claude Opus 5.5 --- system/checks.md | 280 ++++++++++++++++++++----------------------- system/statistics.md | 174 +++++++++++++-------------- 2 files changed, 209 insertions(+), 245 deletions(-) diff --git a/system/checks.md b/system/checks.md index c3af1e9..c7cf151 100644 --- a/system/checks.md +++ b/system/checks.md @@ -2,55 +2,38 @@ **Status: DRAFT** -What a check is, the two checks the rebuild ships, how a check policy is -defined and versioned, how promotion validates against one, and -automatic promotion's design. The [runs](runs) page's Promotion -eligibility paragraph and the [specification](specification)'s -Promotion section state the requirement; this page records how the -rebuild meets it. - -## In plain terms - -A check is a function that looks at one product instance already in -the database and records whether it passes. A check policy is a named, -versioned list of checks with their bounds, some required and some -advisory, and it carries its own approval level: unapproved policies -gate nothing, a trial-approved policy can gate a person's promotion, -and only a team-approved one can gate an automatic one. Promotion runs -the policy's required checks and refuses if any is missing or failed; -nothing about this page changes what promotion already does with a -passing or failing check, it defines what running one means. Automatic -promotion is built and tested, and stays off in production until the -team approves a policy for it. - -Plainly: `rebuild-trial@1`, the policy every promotion in this rebuild -has run under so far, carries a trial approval, not the team's own -sign-off. That trial approval is enough to gate a person's `run promote` -by hand, which is what every demonstration on this page and the -[runs](runs) page has done. It is not enough to gate an automatic -promotion, and no policy the rebuild ships is. The team's own approval -of a policy, and whether any policy the team approves also permits -automatic promotion, are both still open. +A check examines one product instance in the database and records +whether it passes. A check policy names and versions a list of required +and advisory checks with their bounds. Promotion runs the required +checks and refuses if any is missing or failed. This page defines how +checks run without changing how promotion treats their results, meeting +the requirement in the [runs](runs) page's Promotion eligibility +paragraph and the [specification](specification)'s Promotion section. + +Every promotion in the rebuild so far, including every demonstration on +this page and the [runs](runs) page, has used `rebuild-trial@1` for a +person's `run promote` by hand. Its trial approval permits manual +promotion. Automatic promotion is built and tested but stays off in +production: no shipped policy has the team's approval to permit it. ## What a check is A check is a named, versioned Python function over one product -instance, registered in `rapidpipe/checks/`: a registry keyed -`name@version`, declared with `@check("difference-image-statistics", -"1")`. Its signature is `(conn, instance_id, params: dict) -> -CheckResult(outcome 'passed'|'failed', detail dict)`. Running a check -always records a `checks` row: instance, check name, version, the -required flag from the policy that ran it, outcome, and a detail -document of the measurements, the bounds, the params it ran under and -the reason for the outcome; a check that raises records `failed` with -the error in `detail`, never nothing. Recording the params matters -because a policy can override a shipped check's bounds: a result -recorded under other params is not a result under these, so the -promotion gate below only ever counts a row whose recorded params equal -the resolving policy's params for that check. Versions are strings, and -a changed threshold or a changed measurement is a new version, not an -edit to an existing one. A check reads the database; it may read S3 -later, but both checks this page ships are database-only. +instance. It is registered in `rapidpipe/checks/`, keyed by +`name@version`, with a declaration such as +`@check("difference-image-statistics", "1")`. Its signature is +`(conn, instance_id, params: dict) -> CheckResult(outcome 'passed'|'failed', detail dict)`. +Versions are strings; a changed threshold or measurement requires a +new version. Both shipped checks read only the database; future checks +may read S3. + +Running a check always records a `checks` row: instance, check name, +version, the policy's required flag, outcome, and a detail document +containing the measurements, bounds, params and reason for the outcome. +A check that raises records `failed` with the error in `detail`. +Policies can override a shipped check's bounds, so the promotion gate +counts only rows whose recorded params match the resolving policy's +params for that check. ## The two checks @@ -65,24 +48,20 @@ ratio falls within `[ratio_lo, ratio_hi]`. All bounds come from the policy's `params` for that check. `catalog-counts-vs-reference@1` runs over a `source-set` result-set -instance. It now finds its reference by slot rather than walking -dependencies by hand: the candidate's slot already names its exposure, -detector and catalog type ([products](products) page), so the check -looks for another current `source-set` instance sharing that same slot, -any differencer, or, when `params` names a `reference_run`, that run's -instance for the same slot instead. A candidate or a would-be reference -without a slot is not a reference; the check's own runner fills every -candidate's slot first, so a set only misses one when its producer is -still unresolved (below). It compares the candidate's +instance. It finds its reference by slot instead of walking dependencies +by hand. The slot names the candidate's exposure, detector and catalog +type ([products](products) page). The reference is another current +`source-set` instance in the same slot, from any differencer, or, when +`params` names a `reference_run`, that run's instance for the same slot. +A candidate or would-be reference without a slot is not a reference. +The runner fills every candidate's slot first, so a set lacks one only +when its producer is still unresolved. The check compares the candidate's `result_sets.row_count` against the reference's and passes when `|candidate − reference| / reference` is at most `params.tolerance`. When no reference instance exists, the outcome follows `params.missing_reference`, recorded in `detail`; policy `rebuild-trial@1` sets that to pass. -Policy `rebuild-trial@1` marks `difference-image-statistics` required -and `catalog-counts-vs-reference` advisory, required false. - ## Check policies A check policy is a named, versioned TOML file shipped in the package, @@ -91,83 +70,101 @@ same `name@version` key as a check. Its fields are `name`, `version`, `approval` (`none`, `trial` or `team`), `approved_by` (nullable; who granted the recorded approval level), `auto_promote` (boolean), and a `[[checks]]` table per check listing its name, version, kind, required -flag and params. A policy is immutable once landed; a change to any of -these is a new version, never an edit in place. Promotion, manual or -automatic, refuses a policy whose `approval` is `none`; automatic -promotion additionally requires `team` and `auto_promote` true, so a -`trial`-approved policy can gate a person's promotion but never an -automatic one. - -One policy ships with the rebuild. `rebuild-trial@1`'s bounds are set -from the control run's measured values with generous margins: +flag and params. A policy is immutable once landed; changing any field +requires a new version. + +Manual and automatic promotion refuse a policy whose `approval` is +`none`. A `trial`-approved policy can gate a person's promotion; +automatic promotion requires `team` and `auto_promote` true. + +The rebuild ships one policy, `rebuild-trial@1`. It marks +`difference-image-statistics` required and `catalog-counts-vs-reference` +advisory, required false. Its bounds come from the control run's measured +values with generous margins: `scalefacref` in `[1e-3, 1e5]`, `dxrmsfin` and `dyrmsfin` each at most 2.0 pixels, `|dxmedianfin|` and `|dymedianfin|` each at most 1.0 pixel, `nsexcatsources` in `[1000, 1e6]`, and the SExtractor positive-to-negative -ratio in `[0.1, 10]`. `rebuild-strict@1`, the same two checks with -bounds the control run cannot meet (`scalefacref` in `[0.99, 1.01]`, -the same two RMS measurements each at most 0.01 pixel, `nsexcatsources` -at most 1000), is a test fixture that exercises a refusal; it does not -ship, and naming it to a command is an unknown policy. The shipped -policy carries `approval: trial`, `approved_by` recording who -granted it, and `auto_promote` false. This trial approval -satisfies the mechanism the runs and specification pages ask for well -enough to demonstrate manual promotion by hand; it is not the team's own -sign-off, which is a recorded open item, and it does not permit -automatic promotion regardless. There is no policy table in the -database: the promotion row records the policy's version string and the -ids of the check rows it relied on, which is the durable record of what -ran; the TOML file is the versioned definition of what the policy meant -at that version. +ratio in `[0.1, 10]`. It carries `approval: trial`, `approved_by` +recording who granted it, and `auto_promote` false. This trial approval +suffices to demonstrate the manual-promotion mechanism required by the +runs and specification pages; the team's own sign-off remains open. + +`rebuild-strict@1` is a refusal test fixture using the same two checks +with bounds the control run cannot meet: `scalefacref` in +`[0.99, 1.01]`, the same two RMS measurements each at most 0.01 pixel, +and `nsexcatsources` at most 1000. It does not ship; naming it to a +command is an unknown policy. + +There is no policy table in the database. The promotion row durably +records what ran through the policy's version string and the ids of the +check rows it relied on. The TOML file defines what that version meant. + +## The check commands + +`run_policy_checks`, the runner called by both `check run` and the +promotion gate, fills every candidate's slot before running any check. +Thus `catalog-counts-vs-reference` never sees a slot the fill could have +resolved. Checks run on the workstation and reach the database through +the instance role; they do not touch the pipeline image. The [tool](tool) +page gives the full command surface for `check list`, `check run` and +`check show`, including every flag, print format and exit code. ## The promotion gate `promote()` takes an optional check policy. `promote_run` and `run promote` resolve which one to validate against in this order: an explicit `--check-policy`, then the run's own `check_policy_ref`, then -the default `rebuild-trial@1`. Promotion first refuses an unapproved -policy (`approval: none`) before looking at any check result. Every -promotion under an approved policy is then validated under its version: -for each after-instance, every policy check whose kind matches that -instance must have a latest `checks` row, for that check's name, -version and the policy's own params for it, with outcome `passed` when -the policy marks it required; ties in `happened_at` break by row id, and -the read locks the rows it relies on, so a check recorded after the read -started is not counted. A missing or failed required check refuses the -whole promotion, exit 1, naming the instance, the check and the -outcome in the message; a failed or missing advisory check is recorded -but does not refuse. The promotion row records `check_policy_version` -and `check_result_ids`, the rows it relied on. A kind the policy names -no check for passes trivially, and there is no unchecked-exception -flag: a deliverable of a kind the policy does cover always goes through -this gate. - -Promotion also walks every dependency to its roots, not only the -after-instances themselves. `_refuse_unpromoted_ancestor` follows -`dependencies` recursively for each after-instance, cycle-guarded -rather than depth-capped, reading the same instance states `run show` -prints ([runs](runs) page, "Rules"). An ancestor that is itself -`current` or `superseded` passes, since its own chain was judged when -it was promoted, but the walk continues past it rather than stopping -there, so a `candidate` grandparent behind a current parent still -refuses the promotion. An ancestor that is itself an after-instance of -the same promotion request also passes the walk, but gets no exemption -from validation: it is judged in its own right, against the same -eligibility rule and the same check policy every after-instance -answers to. Any ancestor that is neither `current` nor `superseded` nor -a member of the request's own after-instances refuses the whole -promotion, exit 1, naming the after-instance, the ancestor, its kind -and its state: "after instance '01...' depends on '01...' +the default `rebuild-trial@1`. Promotion refuses an unapproved policy +(`approval: none`) before reading any check result. + +For each after-instance, every check in the approved policy whose kind +matches the instance must have a latest `checks` row for its name, +version and policy params. Required checks must have outcome `passed`. +Ties in `happened_at` break by row id. The read locks the rows it relies +on; a check recorded after the read starts is not counted. + +A missing or failed required check refuses the whole promotion, exit 1, +with a message naming the instance, check and outcome. A failed or +missing advisory check is recorded but does not refuse. The promotion +row records `check_policy_version` and `check_result_ids`, the rows it +relied on. A kind with no check in the policy passes trivially. There is +no unchecked-exception flag: a deliverable of a kind covered by the +policy always goes through this gate. + +Acceptance means required checks pass, nothing more ({ref}`acceptance +ruling `). There is no separate acceptance record: +the gate re-reads the latest result of every required check each time. +The `checks` row is the durable record of what let a candidate through; +no judgement is stored on the instance. + +### Dependencies and rollback + +Promotion walks every dependency to its roots. +`_refuse_unpromoted_ancestor` follows `dependencies` recursively for +each after-instance, with a cycle guard and no depth cap. It reads the +same instance states that `run show` prints ([runs](runs) page, +"Rules"). A `current` or `superseded` ancestor passes because its chain +was judged when it was promoted. The walk still continues past it, so a +`candidate` grandparent behind a current parent refuses promotion. + +An ancestor included among the same request's after-instances also +passes the walk, but must independently meet the same eligibility rule +and check policy as every after-instance. Any ancestor that is neither +`current` nor `superseded` nor an after-instance in the request refuses +the whole promotion, exit 1. The message names the after-instance, +ancestor, kind and state: "after instance '01...' depends on '01...' (difference-image, candidate: not current or superseded; promote it first or in the same request); refusing", the hint appearing only for a `candidate` ancestor. A refusal changes no selection or custody; the transaction raises before anything is written. -Rollback skips both gates, the check-policy gate and the ancestor walk: -it restores a selection an earlier promotion already admitted, the same -treatment the released-image check gives it. -Each direct dependency of the restored selection must still be complete, -retained and in project custody; only the walk past those direct -dependencies is skipped. +Rollback skips the check-policy gate and the ancestor walk because it +restores a selection an earlier promotion admitted, the same treatment +the released-image check gives it. Each direct dependency of the +restored selection must still be complete, retained and in project +custody; only the walk past those direct dependencies is skipped. + +### Accepted defect The gate's read of the latest check rows takes its lock `FOR SHARE`; a failed check row committed after that read is not seen by the @@ -175,49 +172,28 @@ promotion it should have refused. The dependency walk reads no check rows: an ancestor's state is custody, completeness and selection. This is an accepted defect, not fixed. -## Acceptance - -Acceptance is required checks pass, nothing more ({ref}`acceptance -ruling `). There is no acceptance record separate -from a `checks` row: the promotion gate above re-reads the latest -result of every required check each time it runs, and that check -result, not a stored judgement on the instance, is the durable record -of what let a candidate through. - ## Automatic promotion -Automatic promotion is designed in and switched off. `run create ---auto-promote [--check-policy P]` is accepted only when the named +`run create --auto-promote [--check-policy P]` is accepted only when the named policy carries the team's approval, recorded as `approval: team` in the policy file, and its `auto_promote` is true; otherwise it is refused, exit 1, with a message naming the policy and stating that team -approval is pending. The shipped policy does not carry the -team's approval, so it does not permit it. At the end -of a `run start` walk, once every unit is complete, the tool calls -`maybe_auto_promote(conn, run_id)`: with the run's `auto_promote` flag +approval is pending. + +At the end of a `run start` walk, once every unit is complete, the tool +calls `maybe_auto_promote(conn, run_id)`: with the run's `auto_promote` flag true, it runs the resolved policy's checks over the run's candidates and promotes on a pass; otherwise it prints `auto-promote off (policy )` and does nothing further, immediately before `start`'s own final `run= state=complete` line. Every run created so far prints -this line, since no shipped policy permits automatic promotion. The -path is exercised end to end against a fixture policy in tests; in -production it can currently only print, because no policy carries the -approval that would let it act. - -## The check commands - -`run_policy_checks`, the runner both `check run` and the promotion gate -call, fills every candidate's slot before running a single check, so -`catalog-counts-vs-reference` never sees a slot the fill itself could -have resolved. Checks run on the workstation, reaching the database -through the instance role; none of this touches the pipeline image. The -full command surface, `check list`, `check run` and `check show`, -including every flag, print format and exit code, is on the -[tool](tool) page. +this line, since no shipped policy permits automatic promotion. +Tests exercise the path end to end with a fixture policy. In production +it can currently only print: no policy carries the approval to act. ## Not decided here - The team's own sign-off on a policy, beyond `rebuild-trial@1`'s trial - approval, and the content of any policy the team approves for - automatic promotion: both scientific, and the team's. + approval, whether an approved policy permits automatic promotion, and + the content of any policy approved for it. These are scientific + questions for the team. - Checks that read S3 rather than only the database. diff --git a/system/statistics.md b/system/statistics.md index 5c10e4e..63d34a4 100644 --- a/system/statistics.md +++ b/system/statistics.md @@ -2,40 +2,37 @@ **Status: DRAFT** -This page covers what the `statistics` stage reads and what it computes. -It also records what lands in `astroobjectsmeta`, the result set the -stage records, its settings and its exit codes. The stage lands on the -pipeline repository's `rebuild` branch: - -- `rapidpipe/stages/statistics.py`; -- `rapidpipe/science/statistics/lightcurve.py`; -- `rapidpipe/settings/statistics.toml`; -- `rapidpipe/db/objects.py`; -- migrations `20260924-03` to `-06`. - -The port is of `dev`'s `pipeline/computeStatisticsForAstroObjects.py`. -Every ported stage minimises differences from `dev` and designs in, -off by default, anything that could be left unused. -The [products](products) page fixes the vocabulary; this page records -how the stage meets it. +`statistics` computes object positions, fluxes and source counts from a +field's association set and records them in `astroobjectsmeta`. The +[products](products) page fixes the vocabulary. + +The stage is ported from `dev`'s +`pipeline/computeStatisticsForAstroObjects.py` onto the pipeline +repository's `rebuild` branch: `rapidpipe/stages/statistics.py`, +`rapidpipe/science/statistics/lightcurve.py`, +`rapidpipe/settings/statistics.toml`, `rapidpipe/db/objects.py`, and +migrations `20260924-03` to `-06`. Every ported stage minimises +differences from `dev` and leaves optional features designed in but off +by default. ## In plain terms `crossmatch` groups a field's sources into objects: each object is a -row of `astroobjects_`, and each `merges_` row links an -object to one of its sources. For one field's association set, the -`statistics` stage gathers every source of every object. From them it -computes the object's mean position and flux, their spread, and how -many sources there are, exactly as `dev` does. It writes one row per -object into `astroobjectsmeta_`. Everything the stage writes for -one association set is one statistics set: a named group of rows that -is either complete or absent. The stage never changes the association -set it reads. +row of `astroobjects_`, and each `merges_` row links it to +one source. `statistics` gathers every source of every object in the +field's association set. It computes each object's mean position and +flux, their spread and the source count, exactly as `dev` does, then +writes one row per object into `astroobjectsmeta_`. These rows +form one statistics set, a named group that is either complete or absent. +The association set stays unchanged. ## Inputs -`--inputs` is `crossmatch`'s completion manifest. The stage reads its -one `association-set` entry: +`--inputs` is `crossmatch`'s completion manifest, though the manifest's +own `stage` need not be `crossmatch`. It must carry exactly one +`association-set` entry for the unit's field. That set must already be +registered; the stage resolves everything else through the database by +instance id: | Field | Used for | |---|---| @@ -44,33 +41,26 @@ one `association-set` entry: | `key.base` | read from the database, not the manifest: the base the set extends | | `key.source_sets` | read from the database, not the manifest: the source sets the set was made from | -The stage does not require the input manifest's own `stage` to be -`crossmatch`, only that it carries exactly one `association-set` entry -for the unit's field. The association set must already be registered: -the stage resolves everything else through the database by instance id. - ### The membership -An association set is **base plus delta**: its membership is its own -rows plus the membership of its base, recursively. The stage takes the chain of instances from -`objects.association_chain`, the input set first, and requires every -set in it to be a complete, retained `association-set`. +An association set contains its own rows plus its base's, recursively: +base plus delta. `objects.association_chain` supplies the chain of +instances, input set first. Every set in the chain must be a complete, +retained `association-set`. -The sources an object's statistics are drawn from are the rows of the -source sets named in the `source_sets` key of every set in the chain. -Each source set resolves to its `sources__` child table by -instance (`objects.source_set_table`). +The statistics use sources from the source sets named in every chain +member's `source_sets` key. `objects.source_set_table` resolves each +source set by instance to its `sources__` child table. -This one rule replaces three steps in `dev`: its `l2files` overlap lookup -of candidate child tables, its `pg_class` check that they exist, and its -`diffimages.vbest > 0` join. An association set names exactly the source -sets it was made from, so "best" is already decided upstream, by the -launcher's input selection. +This membership replaces three steps in `dev`: the `l2files` overlap +lookup of candidate child tables, the `pg_class` existence check and the +`diffimages.vbest > 0` join. The association set names exactly the source +sets it was made from; the launcher's input selection has already +decided which are best. ## What it runs -The stage runs `dev`'s per-field steps, in `dev`'s order, and states -which are ported: +The stage follows `dev`'s per-field order, with these changes: | `dev` step | In the rebuild | |---|---| @@ -83,19 +73,19 @@ which are ported: | one CSV, COPY into `astroobjectsmeta_` | ported, with the run columns and through a temporary table with `ON CONFLICT DO NOTHING` | | drop and recreate `astroobjectsmeta_` before, index and CLUSTER it after, VACUUM ANALYZE or drop it if empty | not ported: the table is made once, with its indexes, and appended to; each attempt writes a new set beside the old ones | -In order, inside one transaction: +The stage runs in one transaction: 1. Resolve the chain, the source sets and their child tables. When the chain names no source sets at all, exit 65; `dev` exits 7 when it finds no source tables. 2. With `[statistics] done_check` on, reuse a complete, retained - statistics set with the same key already written in this run, only - when that set's producing attempt is the calling attempt or an - attempt whose disposition is `succeeded` (the attempt-disposition - join is `db/objects.find_complete_result_set`; the [runs](runs) page - has the general rule). A set left by an attempt that committed rows and - then failed, or by another attempt still without a disposition, is - not reused; a fresh attempt writes its own set instead. + statistics set with the same key already written in this run only if + its producer is the calling attempt or has disposition `succeeded`. + The attempt-disposition join is + `db/objects.find_complete_result_set`; the [runs](runs) page has the + general rule. A fresh attempt writes its own set if the earlier + attempt committed rows then failed, or is another attempt still + without a disposition. 3. Make `astroobjectsmeta_` if it does not exist, through `create_astroobjectsmeta_child_table`. The function applies `dev`'s fillfactor 70, unlogged storage, its `nsources` and Q3C `meanradec` @@ -113,20 +103,20 @@ In order, inside one transaction: 5. Group the rows by `aid`, as `dev` does. An `(aid, sid)` pair that appears under two sets of the chain counts once; the count of such repeats goes in the execution record. -6. Per object, in ascending `aid` order: `compute_radec_statistics` - (`dev`'s `rapid_pipeline_subs.py`, verbatim), `np.mean` and `np.std` - of the fluxes, and the source count. Then one CSV line in `dev`'s - column order, then the run columns. +6. For each object in ascending `aid` order, run + `compute_radec_statistics` (`dev`'s `rapid_pipeline_subs.py`, + verbatim), compute `np.mean` and `np.std` of the fluxes, and count + the sources. Write one CSV line in `dev`'s column order, followed by + the run columns. 7. Register the statistics set, COPY the CSV, read back the row count for the set (it must equal the lines written, else exit 70), and commit. -`compute_radec_statistics` averages the sources' unit vectors, so the -0/360 RA wrap and the poles are handled. The RA spread is the standard -deviation of the shortest signed RA difference scaled by cos(Dec). Its +`compute_radec_statistics` averages the sources' unit vectors to handle +the 0/360 RA wrap and the poles. The RA spread is the standard deviation +of the shortest signed RA difference scaled by cos(Dec). The function's fifth value, the RMS angular spread, is only logged by `dev`: -`astroobjectsmeta` has no column for it. Two consequences, both as in -`dev`: +`astroobjectsmeta` has no column for it. As in `dev`: - A one-source object has standard deviations 0.0, not NaN, because `np.std` is the population standard deviation. @@ -139,10 +129,10 @@ same. ## What lands in `astroobjectsmeta` -One row per object with at least one source in the membership, in -`astroobjectsmeta_`. `dev`'s columns keep their meaning; the last -three are added by migration `20260924-03`, nullable, all set or all -null, and are set on every row the rebuild writes. +`astroobjectsmeta_` holds one row per object with at least one +source in the membership. `dev`'s columns keep their meaning. Migration +`20260924-03` adds the last three: nullable, all set or all null, and set +on every row the rebuild writes. | Column | Source | |---|---| @@ -155,18 +145,16 @@ null, and are set on every row the rebuild writes. | `attempt` | the statistics attempt that wrote the row | | `result_set` | the statistics set's instance | -Keys are set-scoped (products page): `UNIQUE (result_set, aid)`. -Statistics for one object under two sets are two rows. `dev`'s -table-wide primary key on `aid` is not copied onto the rebuild's -per-field tables. +The set-scoped key is `UNIQUE (result_set, aid)` (products page), so +one object's statistics under two sets occupy two rows. The rebuild's +per-field tables do not copy `dev`'s table-wide primary key on `aid`. ## The statistics set -One `statistics-set` result set per attempt, the products page's -database result set. Its unit is the field and its logical key is -`{"membership": }`, which names the exact -membership it describes. The stage writes three things in the same -transaction as the rows: +Each attempt writes one `statistics-set`, the products page's database +result set. Its unit is the field; its logical key, +`{"membership": }`, names the exact membership +it describes. In the same transaction as the rows, the stage records: - its `product_instances` row: kind `statistics-set`, stage `statistics`, this attempt as producer and registrar, custody by run @@ -175,8 +163,8 @@ transaction as the rows: statistics; - a dependency edge to the association set. -An association set whose membership has no `merges` rows gives an empty -set, still complete. +An association set with no `merges` rows in its membership produces an +empty, complete set. The manifest entry has no members; its `registration` block is informational: @@ -193,21 +181,21 @@ source sets resolved. ## Settings -`rapidpipe/settings/statistics.toml`. `dev`'s script reads no -load-bearing setting: its `.ini` reads serve its processing-date scan -or are never used. The rebuild's unit and input replace them. +Settings live in `rapidpipe/settings/statistics.toml`. `dev`'s script +reads no setting that affects the computation: its `.ini` reads serve +the processing-date scan or go unused. The rebuild's unit and input +replace them. | Setting | Default | Meaning | |---|---|---| | `[statistics] done_check` | true | reuse a complete statistics set with the same key already written in this run | | `[statistics] membership` | `association` | what the statistics describe; `pruned` is refused with 64 | -Delivered statistics describe the association set, as `dev` computes -them: `dev` runs crossmatch, then statistics, then prune. The pruned -set as a statistics input is designed in, since the key already names -either kind of set, and left unused ([products](products) page). -The setting names that choice so turning it on later is a settings -change plus the code behind it. +Delivered statistics describe the association set, following `dev`'s +order: crossmatch, statistics, prune. Pruned-set input is designed in +but unused ([products](products) page): the key can already name either +kind of set. Enabling it later needs a settings change and the code +behind it. ## Exit codes @@ -234,8 +222,8 @@ set with the done check off, and base plus delta. ## Open -- `dev`'s aid self-dedupe (`computeStatisticsForAstroObjects.py:214`) - deletes repeated `aid` rows of `astroobjects_`, keeping the - last written. How repeats arose in `dev`, and whether the rebuild's - keep-the-first `ON CONFLICT` choice matches what `dev` intended, - remains a question for Russ. +`dev`'s aid self-dedupe (`computeStatisticsForAstroObjects.py:214`) +deletes repeated `aid` rows of `astroobjects_`, keeping the last +written. How repeats arose in `dev`, and whether the rebuild's +keep-the-first `ON CONFLICT` choice matches what `dev` intended, +remains a question for Russ.