Skip to content

Adopt upstream radiometric calibration QC filter and retrain - #1

Merged
thomasteisberg merged 5 commits into
mainfrom
qc-calibration-filter
Sep 3, 2026
Merged

thomasteisberg merged 5 commits into
mainfrom
qc-calibration-filter

Conversation

@thomasteisberg

@thomasteisberg thomasteisberg commented Sep 2, 2026 •

Copy link
Copy Markdown
Member

Summary

The upstream radar-return-statistics stores now ship per-trace radiometric calibration diagnostics (method 0.4.1: image-combine seam offsets, surface source image, surface ceiling margin). This PR adopts the changelog's suggested QC filter, re-pins all three stores to the calibration snapshots, retrains both required-SNR models, and documents the effect.

  • split.calibration_qc (new, default on): reject traces with a seam step ≥ 3 dB, a surface sampled from a higher-gain image, or a surface within 2 dB of its season's clip level. Rules pass where unmeasured, because the seam check could not run on 92% of 2014_Greenland_P3 and 61% of 2012_Antarctica_DC8. Drops 14% of Antarctic and 13% of Greenland traces after the existing season exclusions.
  • Re-pinned snapshots to the 2026-09-03 method-0.4.1 refresh (antarctica 087JBD7NTAE8BTBTEYSG, greenland ACA8WY61ZF9W6VSBA3HG, ase F6TXWSEQ9RQCD1H7MSMG); upstream verified no pre-existing values changed, so old vs new differs only by the filter. The 0.4.1 refresh over 0.4.0 changed nothing that reaches the model (72 already-rejected 2018 DC8 traces lost an at-ceiling flag); the retrain reproduces the 0.4.0 posteriors exactly.
  • Retrained models (run_ids atten_refl af5514ae3b72, linear c0b385764662): atten_refl CV RMSE 13.02 → 12.87 dB, test 12.72 dB; linear 13.92 → 13.80 dB. Mission tool payload regenerated.
  • Snakefile: each rule depends on a stamp of exactly the config section its run_id hashes (outputs/config_stamps/), so a re-pin or config edit re-runs only the stages it affects.
  • Comparison tooling: scripts/qc_filter_comparison.py, season_crossover_matrix.py --calibration-qc.

Assessment

The filter is sound hygiene with almost no effect on the model. Scored on the same held-out points, pre- and post-QC posteriors differ by 0.03 dB RMSE; every parameter stays within ~2 posterior sd; full-grid predictions shift by −0.2 to +0.6 dB (5th–95th pct). Model-free crossovers show the one real repair: 2016_Greenland_P3 (49% seam rejects) now agrees with the other P3 seasons (offsets ~0 dB, scatter 12 → 8 dB). Season-level offsets (2012 DC8 ~12 dB low, 2017 Basler ~30 dB low, 2013 Greenland 16–19 dB low) are untouched, so exclude_collections stays. Full write-up: agent_notes/20260901-calibration-qc-results.md; docs updated in docs/2_input_data.md and docs/3_model.md.

Test plan

  • uv run pytest -q (68 passed, incl. new tests/unit/test_calibration_qc.py)
  • uv run ruff check src tests scripts
  • uv run snakemake --cores 4 model_all end to end on the re-pinned stores (0 divergences, R̂ ≤ 1.005), twice: 0.4.0 and the 0.4.1 refresh, bit-identical posteriors
  • Section-specific rerun triggers verified with scripted config edits (see commit ec48bfb)

🤖 Generated with Claude Code

https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2

thomasteisberg and others added 4 commits September 2, 2026 13:20
Implements the upstream dataset-changelog (2026-09, calibration method
0.4.0) suggested filtering as split.calibration_qc: reject traces with an
image-combine seam step |img_comb_offset_dB| >= 3 dB, a surface sampled
from a higher-gain image (surface_source_image_index >= 2), or a surface
within 2 dB of its season's fitted clip level. Rules pass where their
input is unmeasured (NaN / unknown): the seam check could not run on 92%
of 2014_Greenland_P3 and 61% of 2012_Antarctica_DC8, so rejecting
unmeasured traces would remove whole seasons. Applied to observations and
non-detections before grid matching; per-rule counts go in the split
manifest. The four calibration columns are carried through augment.

Snakefile: store configs are inputs of the augment rule and model.yaml of
split/train, so a snapshot re-pin or config edit re-runs the right stages
without --forcerun.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
antarctica SYAKG11X8AFFY0H9H65G, greenland JWDABR34HPM816P70FD0, ase
WQCXS05H226PX1KGRZ2G (all "[calibration] saturation second pass",
2026-09-02 00:36 UTC; no pre-existing values changed upstream).

Retrained atten_refl (run_id ef95a4a6fba5) and linear (9b77f5fc2990) on
17,322 training grid points: CV RMSE 13.02 -> 12.87 dB, test RMSE 12.72 dB,
0 divergences. Mission tool payload regenerated. Docs: QC section and
per-rule shares in 2_input_data.md, updated headline numbers and a
"Effect of the radiometric calibration QC" subsection in 3_model.md, all
model figures regenerated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
scripts/qc_filter_comparison.py scores the pre- and post-QC posteriors on
the same held-out points, overlays posteriors in physical units, maps the
prediction difference, and tabulates per-season rejection shares and
threshold sensitivity. season_crossover_matrix.py gains --calibration-qc
(model-free before/after check) and a pairs CSV; posterior_physical.py
exposes physical_panels() for reuse. Results and assessment in
agent_notes/20260901-calibration-qc-results.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
Review found that model.yaml was an input of split and train but not
grid, so a grid-section edit rebuilt downstream stages against a stale
grid while their manifests claimed the new config. Each rule now depends
on a stamp file of exactly the config section its run_id hashes (grid:
inputs+grid; split: inputs+split; train: train+its model entry; augment:
the store's full config), written at parse time only when the hash
changes. First-time stamps get an epoch mtime so existing outputs are not
rebuilt on migration. Verified with dry runs: grid edit -> grid+split+train;
split edit -> split+train; single-model edit -> that model only;
comment-only edits -> nothing; store config edit -> that store's augment
onward.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
@thomasteisberg

thomasteisberg commented Sep 3, 2026 •

Copy link
Copy Markdown
Member Author

Good catch, fixed in ec48bfb (see the commit for details). Rather than adding model.yaml as a grid input (which would also rebuild the slow grid on any train/split edit), each rule now depends on a stamp of exactly the config section its run_id hashes, written by the Snakefile at parse time only when that section's hash changes:

  • grid: inputs + grid
  • split: inputs + split
  • train: train + that model's entry (one stamp per model)
  • augment: the store's full config (replaces the whole-file input, so comment-only edits and re-pin annotations no longer rerun a 10-minute extraction)

First-time stamps get an epoch mtime, so migrating existing outputs does not trigger a rebuild. Verified by dry runs after scripted edits: grid edit → grid + split + train; split edit → split + train (no grid); linear-only entry edit → linear only; draws edit → both models; comment-only edits in either config → nothing; Greenland min_thickness_m edit → Greenland augment onward.

antarctica 087JBD7NTAE8BTBTEYSG, greenland ACA8WY61ZF9W6VSBA3HG, ase
F6TXWSEQ9RQCD1H7MSMG (2026-09-03 15:51 UTC). The 0.4.1 refresh drops
2018_Antarctica_DC8's rising img2 envelope as a ceiling, backfills 49
Antarctic frames' seam checks and adds variable attrs; the only change
reaching the model is 72 2018 DC8 traces losing an at-ceiling flag, all
already rejected by the img2 rule. The split is identical apart from
run_ids and both models retrain to bit-identical posteriors (run_ids:
split 4ff998437f3b, atten_refl af5514ae3b72, linear c0b385764662). The
config stamps triggered the augment reruns without --forcerun. Mission
tool payload regenerated; docs/notes updated to 0.4.1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y73QhGDAtbDwo9ZnJ2N4F2
@thomasteisberg

Copy link
Copy Markdown
Member Author

Re-pinned to the 0.4.1 refresh (2026-09-03 15:51 UTC snapshots) and retrained in 6a1fb54. The refresh's only model-visible change is 72 traces of 2018_Antarctica_DC8 losing their at-ceiling flag (its rising img2 envelope no longer counts as a ceiling), and all 72 were already rejected by the img2 rule, so the split is identical apart from run_ids and both models retrain to bit-identical posteriors. Only provenance (snapshot IDs, run_ids, mission-tool meta) and the doc references changed; the figures did not need regenerating. The config stamps triggered both augment reruns on their own.

@thomasteisberg
thomasteisberg merged commit 6016a49 into main Sep 3, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant