Adopt upstream radiometric calibration QC filter and retrain - #1
Conversation
Implements the upstream dataset-changelog (2026-09, calibration method 0.4.0) suggested filtering as split.calibration_qc: reject traces with an image-combine seam step |img_comb_offset_dB| >= 3 dB, a surface sampled from a higher-gain image (surface_source_image_index >= 2), or a surface within 2 dB of its season's fitted clip level. Rules pass where their input is unmeasured (NaN / unknown): the seam check could not run on 92% of 2014_Greenland_P3 and 61% of 2012_Antarctica_DC8, so rejecting unmeasured traces would remove whole seasons. Applied to observations and non-detections before grid matching; per-rule counts go in the split manifest. The four calibration columns are carried through augment. Snakefile: store configs are inputs of the augment rule and model.yaml of split/train, so a snapshot re-pin or config edit re-runs the right stages without --forcerun. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
antarctica SYAKG11X8AFFY0H9H65G, greenland JWDABR34HPM816P70FD0, ase WQCXS05H226PX1KGRZ2G (all "[calibration] saturation second pass", 2026-09-02 00:36 UTC; no pre-existing values changed upstream). Retrained atten_refl (run_id ef95a4a6fba5) and linear (9b77f5fc2990) on 17,322 training grid points: CV RMSE 13.02 -> 12.87 dB, test RMSE 12.72 dB, 0 divergences. Mission tool payload regenerated. Docs: QC section and per-rule shares in 2_input_data.md, updated headline numbers and a "Effect of the radiometric calibration QC" subsection in 3_model.md, all model figures regenerated. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
scripts/qc_filter_comparison.py scores the pre- and post-QC posteriors on the same held-out points, overlays posteriors in physical units, maps the prediction difference, and tabulates per-season rejection shares and threshold sensitivity. season_crossover_matrix.py gains --calibration-qc (model-free before/after check) and a pairs CSV; posterior_physical.py exposes physical_panels() for reuse. Results and assessment in agent_notes/20260901-calibration-qc-results.md. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
Review found that model.yaml was an input of split and train but not grid, so a grid-section edit rebuilt downstream stages against a stale grid while their manifests claimed the new config. Each rule now depends on a stamp file of exactly the config section its run_id hashes (grid: inputs+grid; split: inputs+split; train: train+its model entry; augment: the store's full config), written at parse time only when the hash changes. First-time stamps get an epoch mtime so existing outputs are not rebuilt on migration. Verified with dry runs: grid edit -> grid+split+train; split edit -> split+train; single-model edit -> that model only; comment-only edits -> nothing; store config edit -> that store's augment onward. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
|
Good catch, fixed in ec48bfb (see the commit for details). Rather than adding
First-time stamps get an epoch mtime, so migrating existing outputs does not trigger a rebuild. Verified by dry runs after scripted edits: grid edit → grid + split + train; split edit → split + train (no grid); linear-only entry edit → linear only; |
antarctica 087JBD7NTAE8BTBTEYSG, greenland ACA8WY61ZF9W6VSBA3HG, ase F6TXWSEQ9RQCD1H7MSMG (2026-09-03 15:51 UTC). The 0.4.1 refresh drops 2018_Antarctica_DC8's rising img2 envelope as a ceiling, backfills 49 Antarctic frames' seam checks and adds variable attrs; the only change reaching the model is 72 2018 DC8 traces losing an at-ceiling flag, all already rejected by the img2 rule. The split is identical apart from run_ids and both models retrain to bit-identical posteriors (run_ids: split 4ff998437f3b, atten_refl af5514ae3b72, linear c0b385764662). The config stamps triggered the augment reruns without --forcerun. Mission tool payload regenerated; docs/notes updated to 0.4.1. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
|
Re-pinned to the 0.4.1 refresh (2026-09-03 15:51 UTC snapshots) and retrained in 6a1fb54. The refresh's only model-visible change is 72 traces of 2018_Antarctica_DC8 losing their at-ceiling flag (its rising img2 envelope no longer counts as a ceiling), and all 72 were already rejected by the img2 rule, so the split is identical apart from run_ids and both models retrain to bit-identical posteriors. Only provenance (snapshot IDs, run_ids, mission-tool meta) and the doc references changed; the figures did not need regenerating. The config stamps triggered both augment reruns on their own. |
Summary
The upstream
radar-return-statisticsstores now ship per-trace radiometric calibration diagnostics (method 0.4.1: image-combine seam offsets, surface source image, surface ceiling margin). This PR adopts the changelog's suggested QC filter, re-pins all three stores to the calibration snapshots, retrains both required-SNR models, and documents the effect.split.calibration_qc(new, default on): reject traces with a seam step ≥ 3 dB, a surface sampled from a higher-gain image, or a surface within 2 dB of its season's clip level. Rules pass where unmeasured, because the seam check could not run on 92% of 2014_Greenland_P3 and 61% of 2012_Antarctica_DC8. Drops 14% of Antarctic and 13% of Greenland traces after the existing season exclusions.087JBD7NTAE8BTBTEYSG, greenlandACA8WY61ZF9W6VSBA3HG, aseF6TXWSEQ9RQCD1H7MSMG); upstream verified no pre-existing values changed, so old vs new differs only by the filter. The 0.4.1 refresh over 0.4.0 changed nothing that reaches the model (72 already-rejected 2018 DC8 traces lost an at-ceiling flag); the retrain reproduces the 0.4.0 posteriors exactly.af5514ae3b72, linearc0b385764662): atten_refl CV RMSE 13.02 → 12.87 dB, test 12.72 dB; linear 13.92 → 13.80 dB. Mission tool payload regenerated.outputs/config_stamps/), so a re-pin or config edit re-runs only the stages it affects.scripts/qc_filter_comparison.py,season_crossover_matrix.py --calibration-qc.Assessment
The filter is sound hygiene with almost no effect on the model. Scored on the same held-out points, pre- and post-QC posteriors differ by 0.03 dB RMSE; every parameter stays within ~2 posterior sd; full-grid predictions shift by −0.2 to +0.6 dB (5th–95th pct). Model-free crossovers show the one real repair: 2016_Greenland_P3 (49% seam rejects) now agrees with the other P3 seasons (offsets ~0 dB, scatter 12 → 8 dB). Season-level offsets (2012 DC8 ~12 dB low, 2017 Basler ~30 dB low, 2013 Greenland 16–19 dB low) are untouched, so
exclude_collectionsstays. Full write-up:agent_notes/20260901-calibration-qc-results.md; docs updated indocs/2_input_data.mdanddocs/3_model.md.Test plan
uv run pytest -q(68 passed, incl. newtests/unit/test_calibration_qc.py)uv run ruff check src tests scriptsuv run snakemake --cores 4 model_allend to end on the re-pinned stores (0 divergences, R̂ ≤ 1.005), twice: 0.4.0 and the 0.4.1 refresh, bit-identical posteriors🤖 Generated with Claude Code
https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2