diff --git a/docs/2_input_data.md b/docs/2_input_data.md index ec98ecf..50d4d42 100644 --- a/docs/2_input_data.md +++ b/docs/2_input_data.md @@ -26,7 +26,7 @@ Traces with radar-derived thickness under 100 m are dropped. (This is configured ### Radiometric calibration QC -Two instrument effects can bias RSSNR at the trace level, and since 2026-09 the upstream stores ship per-trace diagnostics for both (calibration method 0.4.1; see the [dataset changelog](https://github.com/englacial/radar-return-statistics/blob/main/docs/dataset_changelog.md)): +Two instrument effects can bias RSSNR at the trace level, and since 2026-09-03 the upstream stores include per-trace diagnostics for both (calibration method 0.4.1; see the [dataset changelog](https://github.com/englacial/radar-return-statistics/blob/main/docs/dataset_changelog.md)): * **Image-combine seam steps.** The MCoRDS products stitch a low-gain image (surface) onto higher-gain images (deep ice/bed). A miscalibrated stitch puts a power step at the seam, which biases bed power relative to surface power by that step. `img_comb_offset_dB` is the residual step measured in each frame's individual images. * **Surface saturation.** Where the surface return clips the receiver, surface power is underestimated and RSSNR is biased low. `surface_ceiling_margin_dB` is the distance below the season's fitted clip level, and `surface_source_image_index` records whether the surface sample came from the low-gain image (1) or a higher-gain image (≥2, likely saturated with a season-dependent low bias). @@ -42,11 +42,6 @@ The split stage applies the upstream-suggested filter (`split.calibration_qc` in (Shares are of the traces entering the split stage, i.e. after the season exclusions above; a trace failing several rules is counted once, in the first row it fails.) -Every rule **passes where its input is unmeasured** (NaN offset because the seam check could not run — no published images, insufficient overlap; NaN margin because no credible ceiling was fitted; unknown source image). That is the changelog's "no evidence either way" reading, and it matters: the seam check could not run on 92% of 2014_Greenland_P3 and 61% of 2012_Antarctica_DC8 at their flight geometry, so dropping unmeasured traces would remove whole seasons (35% of Antarctic and 50% of Greenland traces) rather than bad traces. The img2 rule is kept as suggested because img2-sourced traces have median RSSNR 8–17 dB below img1 traces of the same DC8 season, consistent with the documented surface-power low bias. - -The filter is a trace-level cleanup, not a season-level recalibration. Retraining with it (2026-09-01, `outputs/qc_filter/`, `uv run python scripts/qc_filter_comparison.py`) lowers the atten_refl spatial-CV RMSE from 13.02 to 12.87 dB, leaves held-out test RMSE within 0.05 dB when both posteriors are scored on the same points, and shifts full-grid predictions by less than 1 dB (5th–95th percentile −0.3 to +0.6 dB). The crossover matrices above are almost unchanged by it (`scripts/season_crossover_matrix.py --calibration-qc`): within-season scatter tightens for the seam-affected seasons, but the season-to-season offsets (e.g. 2012_Antarctica_DC8 ~12 dB below the later DC8 seasons) remain, and the QC alone catches only 2% of 2013_Greenland_P3 and 46% of 2017_Antarctica_Basler — so the two season exclusions stay. - - ## Do the surveys agree with each other? The full dataset represents more than a decade of data spanning multiple radar instrument generations and two different lineages (CReSIS's MCoRDS and UTIG's HiCARS/MARFA radars). Before pooling this data, it's worth asking whether they measure the same thing. Wherever these flight lines cross (defined as within 500 m here), we can look for biases between seasons. @@ -71,8 +66,6 @@ cp outputs/model/analysis/season_crossover_matrix_*.png docs/figures/ There are two notable outlier seasons: `2017_Antarctica_Basler` and `2013_Greenland_P3`. These two seasons show significant stitching artifacts that likely explain these offsets. They are excluded from further analysis for now. -> **Stitching artifacts:** Both the MCoRDS and HiCARS/MARFA systems rely on some sort of stitching across either multiple distinct waveforms or ADC channels to achieve their roughly ~100-120 dB dynamic range. If this stitching is incorrectly calibrated, it can artificially make the bed return stronger or weaker relative to the surface than it actually is. More effectively detecting and filtering this out is an ongoing area of work. - The good news is that most seasons agree with each other reasonably well, even across the two institutions. Jumping ahead a little bit (this part depends on a model to train), another way to look at this is to compare against out-of-fold residuals. (Note that folds are spatially blocked. More on this in the next section.) diff --git a/docs/3_model.md b/docs/3_model.md index 114fd71..366abdb 100644 --- a/docs/3_model.md +++ b/docs/3_model.md @@ -14,7 +14,7 @@ $$ \mu_i = a_i H_i - r_i $$ -Where $H_i$ is the ice thickness, $a_i$ can be interpreted as an attenuation rate, and $r_i$ can be interpreted as a basal reflectivity. $a_i$ and $r_i$ are modelled as linear functions of covariates $\mathbf{x}_i = [\,T_{\text{air}},\ \text{speed},\ \text{GHF}\,]_i$, indicators $g_i$ (Greenland) and $f_i$ (floating). +Where $H_i$ is the ice thickness, $a_i$ can be interpreted as an attenuation rate, and $r_i$ can be interpreted as a basal reflectivity. $a_i$ and $r_i$ are modelled as linear functions of covariates $`\mathbf{x}_i = [\,T_{\text{air}},\ \text{speed},\ \text{GHF}\,]_i`$, indicators $g_i$ (Greenland) and $f_i$ (floating). $$\mathbf{c}_i = \bigl[\, \mathbf{x}_i,\ g_i,\ f_i,\ g_i\mathbf{x}_i \,\bigr]$$ @@ -127,7 +127,9 @@ uv run python scripts/posterior_physical.py cp outputs/model/analysis/posterior_physical.png docs/figures/ ``` -The headline accuracy and calibration numbers (trained 2026-09-03 on the method-0.4.1 calibration snapshots with the radiometric calibration QC filter described in [Input data](2_input_data.md#radiometric-calibration-qc); 17,322 training grid points, 1,171 of them censored, plus 187 non-detections): +Key metrics are reported below for the last model training: + +**Last updated:** 2026-09-03 on the method-0.4.1 calibration snapshots with the radiometric calibration QC filter | quantity | value | |---|---| @@ -138,10 +140,16 @@ The headline accuracy and calibration numbers (trained 2026-09-03 on the method- | Fully-linear baseline (same layers) | CV 13.80 dB / test 13.98 dB | | Sampler diagnostics | 0 divergences, R̂ ≤ 1.005 | -Posterior distributions of all 21 learned parameters, converted to physical units (the z-score normalization is an invertible affine transform, and the normalizer constants are stored in `posterior.nc`, so this conversion is exact). Attenuation-side parameters become two-way dB/km via σ_target/σ_thickness; reflectivity-side parameters become dB contributions to RSSNR (sign-flipped for the −refl convention); covariate effects are fully per-unit (e.g. dB/km/K, dB/km/(mW/m²)); θ, τ, and σ are natively in dB. Intercept-like values are referenced to the mean covariate conditions of the training set. +Posterior distributions of all 21 learned parameters, converted to physical units (normalizer constants are stored in `posterior.nc`) are shown below. +Note that: +* Attenuation-side parameters become two-way dB/km via σ_target/σ_thickness +* Reflectivity-side parameters become dB contributions to RSSNR (sign-flipped for the `−refl` convention) +* Covariate effects are fully per-unit (e.g. dB/km/K, dB/km/(mW/m²)) +* θ, τ, and σ are natively in dB +* Intercept-like values are referenced to the mean covariate conditions of the training set.  -*Posteriors in physical units, with the headline CV and held-out test RMSE. The 11.6 dB/km one-way (23.2 two-way) depth-averaged attenuation rate at mean conditions falls in the physically expected range. Posterior widths are small because n ≈ 17k; the meaningful uncertainty is the 12.7 dB residual σ. The 0/1 indicators are reported as the step between their states; the interaction panels are Greenland's *offset* from the Antarctic slope, not an absolute slope.* +*Posteriors in physical units, with the headline CV and held-out test RMSE. The 11.6 dB/km one-way (23.2 two-way) depth-averaged attenuation rate at mean conditions falls in the physically expected range. Posterior widths are small; the meaningful uncertainty is the 12.7 dB residual σ. The 0/1 indicators are reported as the step between their states; the interaction panels are Greenland's offset from the Antarctic slope, not an absolute slope.* The distribution of observed and posterior predicted RSSNR values are shown below by ice sheet: @@ -169,12 +177,12 @@ Maps of the mean and 80th percentile predictions for both ice sheets are shown b ### Effect of the radiometric calibration QC -The model above is the first trained after the upstream stores shipped per-trace seam-step and surface-saturation diagnostics, with the suggested filter applied (`split.calibration_qc`). The filter drops about 14% of Antarctic and 13% of Greenland traces but leaves the fit essentially unchanged: scored on the same held-out points, the pre- and post-QC posteriors differ by 0.03 dB in RMSE, every parameter stays within about two posterior standard deviations of its previous value, and the predicted maps move by less than 1 dB almost everywhere. +As of 2026-09-03, the input data is filtered using QC flags that identify surface saturation and image combine errors that could lead to radiometric miscalibration. The filter drops about 14% of Antarctic and 13% of Greenland traces but leaves the model fit essentially unchanged. Scored on the same held-out points, the pre- and post-QC posteriors differ by 0.03 dB in RMSE and the predicted maps move by less than 1 dB almost everywhere.  *Change in posterior-mean required surface SNR from adopting the calibration QC filter (new − previous model). Median +0.01 dB (Antarctica) and +0.07 dB (Greenland); 5th–95th percentile within ±0.6 dB.* -The season-level offsets visible in the crossover matrices and in the out-of-fold residuals by season are not touched by a trace-level filter and remain the largest known calibration issue. To reproduce the comparison (requires a copy of the previous `outputs/model` and augment parquets at `outputs/baseline_20260807/`): +To reproduce the comparison (requires a copy of the previous `outputs/model` and augment parquets at `outputs/baseline_20260807/`): ``` uv run python scripts/qc_filter_comparison.py diff --git a/mission_design_tool/index.html b/mission_design_tool/index.html index 3c2d710..acfdcdc 100644 --- a/mission_design_tool/index.html +++ b/mission_design_tool/index.html @@ -30,6 +30,9 @@
This tool is designed to build intuition about nadir-pointing ice-penetrating radar sounder instruments on novel platforms. Start from a preset below or enter custom diff --git a/mission_design_tool/style.css b/mission_design_tool/style.css index a6e3e6c..cc44c47 100644 --- a/mission_design_tool/style.css +++ b/mission_design_tool/style.css @@ -147,6 +147,8 @@ h3 { font-size: 14px; margin: 0 0 10px; } .intro { margin-top: 28px; } .intro p { margin: 0 0 14px; } .lede { font-size: 17px; line-height: 1.5; } +.authors { color: var(--muted); font-size: 12.5px; line-height: 1.5; margin: 0 0 10px; } +.authors .credit-note { white-space: nowrap; } .note { color: var(--muted); font-size: 13.5px; } /* Equations: pre.eq is the display block, code.eq the inline form. `code` is inline by default — a `pre` used mid-sentence is block-level and makes the