From dd8ac2a9addf6a05015a4dac34ca179574cc6592 Mon Sep 17 00:00:00 2001 From: Ben Rusholme Date: Mon, 28 Sep 2026 11:38:56 -0700 Subject: [PATCH] openuniv_eval, pipeline_evaluation_metrics, archive: clarity pass (Codex) Co-Authored-By: Claude Opus 5.5 --- .../openuniv_eval.rst | 71 ++-- .../pipeline_evaluation_metrics.rst | 396 ++++++++++-------- docs/source/archive/archive.rst | 107 +++-- 3 files changed, 307 insertions(+), 267 deletions(-) diff --git a/docs/source/analyses/pipeline_evaluation_metrics/openuniv_eval.rst b/docs/source/analyses/pipeline_evaluation_metrics/openuniv_eval.rst index bf54d273..c6e4ce06 100644 --- a/docs/source/analyses/pipeline_evaluation_metrics/openuniv_eval.rst +++ b/docs/source/analyses/pipeline_evaluation_metrics/openuniv_eval.rst @@ -1,30 +1,29 @@ Evaluation of OpenUniverse Simulated Data Pipeline Tests ################################################################## -Overview -************************************ - -The test evalutions described below are organized by processing date. See :ref:`testing` for specifications -of each test run. +Evaluations are organized by processing date. See :ref:`testing` for each +test run's specifications. 9/27/2025 ************************************ -This test run consists of 6,875 science images across 7 filters. We evalutate the performance of the pipeline -for two difference image-subtraction methods: ZOGY and SFFT (with PSF cross-convolution). - +This run evaluates pipeline performance on 6,875 science images across +7 filters using two difference image-subtraction methods: ZOGY and SFFT +(with PSF cross-convolution). Source Injections ==================================== -Fake-source injections were performed with the scheme described in :ref:`fake_source_injection` with -100 sources per image and ``size_factor`` :math:`= 1.5`. +Fake sources were injected as described in :ref:`fake_source_injection`, +with 100 sources per image and ``size_factor`` :math:`= 1.5`. Source Detection and Filtering ==================================== -We evalute the performance of source detection with SExtractor, with detection performed on the matched-filtered difference images -where available (Scorr for ZOGY and cross-convolved difference images for SFFT). -SExtractor photometry is performed on the difference images. The primary SExtractor parameters used for source detection were: + +SExtractor detects sources on matched-filtered difference images where +available (Scorr for ZOGY and cross-convolved difference images for SFFT) +and measures photometry on the difference images. The primary detection +parameters were: * ``DETECT_MINAREA`` :math:`=5` * ``DETECT_THRESH`` :math:`=2.5` @@ -34,7 +33,8 @@ SExtractor photometry is performed on the difference images. The primary SExtrac * ``DEBLEND_NTHRESH`` :math:`=32` * ``PHOT_APERTURES`` :math:`=2.0,3.0,4.0,6.0,10.0,14.0` (aperture diameter in pixels) -Raw catalogs were filtered as described in :ref:`filtering` with the following thresholds: +Raw catalogs were filtered as described in :ref:`filtering`, using these +thresholds: 1. Catalog level a. ``diffimedgetol`` :math:`=5` @@ -44,12 +44,13 @@ Raw catalogs were filtered as described in :ref:`filtering` with the following t 2. Pixel-value level a. ``nnegthres`` :math:`=18` b. ``nbadthres`` :math:`=12` - c. ``sumratthres`` :math:`=0.25` -3. PSF photometry + c. ``sumratthres`` :math:`=0.25` +3. PSF photometry a. ``rchipsfthres`` :math:`=10` b. ``magdiffthres`` :math:`=0.3` -Detections are cross-matched to the injected source catalogs using a matching radius of ``injmatchrad`` :math:`=0.5` WFI pixels. +Detections are cross-matched to injected-source catalogs within +``injmatchrad`` :math:`=0.5` WFI pixels. Filtering results: @@ -68,14 +69,28 @@ SFFT H158 1142 2935660 74690 148362 (5.1%) 66434 (8 Evaluation of Figures of Merit ==================================== -We evaluate the Figure of Merit (FOM; as defined in :ref:`figure_of_merit`) for both difference imaging methods for H158. We define two -groups sources for this analysis based on their separations from the core of the nearest truth-catalog galaxy brighter than -``galmatchthres`` :math:`= 25` mag, off-nuclear sources at :math:`\gt 1.5 \times` ``injmatchrad`` and nuclear sources -at :math:`\leq 1.5 \times` ``injmatchrad``. This grouping allows us to account for possible contamination of the recovered True Positive (TP) -sample by spurious detections arising from imperfect galaxy subtraction in the nuclear regions. We apply overall relative weights of 1.0 and 0.5 -for the off-nuclear and nuclear source groups, respectively, in the FOM calculations. The relative weights for each term in the FOM -are :math:`w_{\mathrm{th}} = 1.0`, :math:`w_{80} = 0.8`, :math:`w_{20} = 0.6`, :math:`w_{5\sigma} = 0.3`, and :math:`w_{\mathrm{ph}10} = 0.2`. -We adopt ``fpratetol`` :math:`=10` as the maximum acceptable false-positive rate per image. + +The Figure of Merit (FOM; defined in :ref:`figure_of_merit`) is evaluated +for both methods in H158. Sources are grouped by their separation from +the core of the nearest truth-catalog galaxy brighter than +``galmatchthres`` :math:`= 25` mag: + +* Off-nuclear: :math:`\gt 1.5 \times` ``injmatchrad``. +* Nuclear: :math:`\leq 1.5 \times` ``injmatchrad``. + +This grouping accounts for possible contamination of the recovered True +Positive (TP) sample by spurious detections from imperfect galaxy +subtraction in nuclear regions. The FOM calculations assign overall +relative weights of 1.0 to off-nuclear sources and 0.5 to nuclear sources. +The relative weights of the FOM terms are :math:`w_{\mathrm{th}} = 1.0`, +:math:`w_{80} = 0.8`, :math:`w_{20} = 0.6`, :math:`w_{5\sigma} = 0.3`, and +:math:`w_{\mathrm{ph}10} = 0.2`. The maximum acceptable false-positive +rate per image is ``fpratetol`` :math:`=10`. + +.. note:: + This preliminary analysis illustrates how to evaluate RAPID pipeline + performance using the FOM. Thresholds, weights, and grouping criteria + for uses such as algorithm down-selection remain under development. ====== ====== ======== ======================= ============== ============== =================== ========================= ========== Method Filter Group :math:`m_{\mathrm{th}}` :math:`m_{80}` :math:`m_{20}` :math:`m_{5\sigma}` :math:`m_{\mathrm{ph}10}` FOM @@ -88,14 +103,8 @@ ZOGY H158 Overall 24.60 23.95 25. SFFT H158 Overall 25.59 24.66 25.59 26.78 24.49 **25.38** ====== ====== ======== ======================= ============== ============== =================== ========================= ========== -.. note:: - This analysis is preliminary and meant to illustrate the procedure for evaluation the performance of the RAPID pipeline - based on the FOM. The chosen thresholds, weights, and grouping criteria that will be used for, e.g., algorithm down-selection, - are under development. - .. toctree:: :maxdepth: 1 FOM_20250927_ZOGY_H158.rst FOM_20250927_SFFT_H158.rst - diff --git a/docs/source/analyses/pipeline_evaluation_metrics/pipeline_evaluation_metrics.rst b/docs/source/analyses/pipeline_evaluation_metrics/pipeline_evaluation_metrics.rst index dd1ae2e0..28853970 100644 --- a/docs/source/analyses/pipeline_evaluation_metrics/pipeline_evaluation_metrics.rst +++ b/docs/source/analyses/pipeline_evaluation_metrics/pipeline_evaluation_metrics.rst @@ -1,131 +1,143 @@ RAPID Pipeline Evaluation -#################################################### +######################### Overview -************************************ +******** -Here we describe the procedures for evaluating the performance of the RAPID pipeline, -downselecting algorithms (e.g., for difference imaging and source detection), and tuning -parameters to optimize performance. This procedure is in development, and will be refined -as we continue to develop the pipeline and evaluate its performance on simulated data sets. -In preparation for Roman launch, we will lay out specific plans and timelines to evaluate -the pipeline and tune parameters with in-flight data. These plans will be developed based on -scheduled Roman observations as they become available. +These procedures evaluate RAPID pipeline performance, guide algorithm +down-selection (e.g., for difference imaging and source detection), and +support parameter tuning. They remain under development as the pipeline +evolves and is evaluated on simulated data sets. Before Roman launch, we +will develop specific plans and timelines for evaluation and tuning with +in-flight data, based on scheduled Roman observations as they become +available. .. _fake_source_injection: -Source Injections -********************************************** -Evalution of the RAPID pipeline performance relies on the injection of synthetic point sources. -The currently implemented approach described below was developed for injection into OpenUniverse simulated images -of a nominal implementation of the HLTDS. - -* Injections associated with detected sources/galaxies - - * Simple source detection and deblending is performed with **PhotUtils** on the science image with - a detection threshold of :math:`10\sigma` above the median background level. The bulk of these detections - will be galaxies in the field, but some may be stars or associated with the wings or diffraction spikes of - bright stars. - - * A random subset of the detected sources are selected for injections based on the desired number of injections - to perform per image, :math:`N_{\mathrm{inj}}`. - - * The location of each injection is randomly offset from the detected source centroid in both x and y using a - uniform distribution of half-width set to the detected source's estimated semi-major axis (``semimajor_sigma`` - measured by **PhotUtils**) multiplied by a scaling factor, ``size_factor``. - - * The flux in image counts (electrons) of each injection is drawn from a uniform distrubtion of magnitdues between - 21 and 28 AB mag, using the appropiate zeropoint for the image filter, Roman SCA, and exposure time. - - * A PSF for each injection is computed using the **Roman I-Sim** tool **romanisim.image.make_one_psf** for the given filter, SCA, - and detector position, based on the **galsim.roman** module, to emulate the PSFs used for the original OpenUniverse - Roman simulations. Chromatic effects are currently ignored. - - * The injection is added to the science image with **romanisim.image.add_objects_to_image** at the specified location, - with the appropriate PSF and flux and including poisson noise. - -Additional schemes for source injections currently under consideration for implementation include: - -* Injections at random positions in the image or a pre-defined grid, independent of detected sources. This scheme would be - useful to evaluate the performance of the pipeline for hostless transients, or in crowded fields (e.g., for the GBTDS and - associated simulations). -* Injections in time-series images of a given field at specified sky locations with pre-defined light curves. This could also - include the injection of a stellar counterparts in the images used to generate reference mosaics to evaluate the performance - of the pipeline for variable stars. +Source Injections +***************** + +Evaluation relies on synthetic point-source injections. The implemented +scheme associates injections with detected sources/galaxies and was +developed for OpenUniverse simulated images of a nominal HLTDS +implementation: + +* Perform simple source detection and deblending on the science image with + **PhotUtils**, at a threshold of :math:`10\sigma` above the median + background. Most detections will be galaxies, but some may be stars or + features in bright stars' wings or diffraction spikes. +* Select a random subset of detections based on the desired number of + injections per image, :math:`N_{\mathrm{inj}}`. +* Offset each injection from the detected source centroid in both x and y, + using a uniform distribution whose half-width is the estimated semi-major + axis (``semimajor_sigma`` measured by **PhotUtils**) multiplied by + ``size_factor``. +* Draw magnitudes uniformly between 21 and 28 AB mag and convert to image + counts (electrons) using the appropriate zeropoint for the image filter, + Roman SCA, and exposure time. +* Compute each PSF with the **Roman I-Sim** tool + **romanisim.image.make_one_psf** for the filter, SCA, and detector + position. This uses **galsim.roman** to emulate the original OpenUniverse + Roman simulation PSFs; chromatic effects are currently ignored. +* Add each injection to the science image with + **romanisim.image.add_objects_to_image**, using the specified location, + PSF, and flux, including Poisson noise. + +Additional schemes under consideration are: + +* Random positions or a pre-defined grid, independent of detected sources, + to evaluate hostless transients or crowded fields (e.g., the GBTDS and + associated simulations). +* Time-series injections at specified sky locations in a given field, with + pre-defined light curves. Stellar counterparts could also be injected + into images used for reference mosaics to evaluate variable stars. .. _filtering: Detection Catalog Filtering -********************************************************** +*************************** -Raw detection catalogs generated by the RAPID pipeline will contain a significant number of spurious detections arising from -imperfect image subtraction resulting in significant residuals from static galaxies and stellar sources, noise fluctuations, -and detector artifacts such as unflagged hot-pixels or cosmic-ray hits. A series of filtering steps, in order of -increasing computational expense, are performed to clean the raw catalogs and provide metrics/features as inputs to machine-learning -based real-bogus (RB) classifier. This procedure adapted from that developed for the `ZTF Science Data System`_, and is under active -development. +Raw RAPID detection catalogs will contain many spurious detections from +imperfect subtraction of static galaxies and stars, noise fluctuations, +and detector artifacts such as unflagged hot pixels or cosmic-ray hits. +The following filters, ordered by increasing computational expense, clean +the catalogs and supply metrics/features to a machine-learning real-bogus +(RB) classifier. Adapted from the `ZTF Science Data System`_, this +procedure remains under active development. .. _ZTF Science Data System: https://irsa.ipac.caltech.edu/data/ZTF/docs/ztf_explanatory_supplement.pdf -1. Catalog-level filtering based on measurements performed during source detection (e.g., by **SExtractor** or **Photutils DAOStarFinder**): - - a. ``mindedge`` :math:`\gt` ``diffimedgetol``: Minimum distance from any image edge in pixels. - - b. ``snrap3pix`` :math:`\geq` ``snrthres``: Signal-to-noise ratio (S/N) measured in a 3-pix diameter aperature at the candidate position - in the difference image and corresponding uncertainty map. A initial threshold of ``snrthres`` :math:`=5` is adopted. - - c. ``elong`` :math:`\leq` ``elongthres``: Source elongation (ratio A/B of semi-major to semi-minor axis). - - d. (``apfluxratio`` :math:`\geq` ``apfluxratiothreslow``) and (``apfluxratio`` :math:`\leq` ``apfluxratiothreshigh``): Ratio of flux - measured in a 3-pixel diameter aperture to that measured in a 6-pixel diameter aperture. - - .. note:: - These filtering steps are based on measurements available in SExtractor catalogs. If **Photutils DAOStarFinder** or another detection - method is employed, analagous measurements will be used. Additional metrics like ``sharpness`` and ``roundness`` estimated by DAOStarFinder - could also be employed. +1. Catalog-level measurements from source detection (e.g., by + **SExtractor** or **Photutils DAOStarFinder**): + + a. ``mindedge`` :math:`\gt` ``diffimedgetol``: Minimum distance from any + image edge, in pixels. + b. ``snrap3pix`` :math:`\geq` ``snrthres``: Signal-to-noise ratio (S/N) + in a 3-pix diameter aperture at the candidate position, measured from + the difference image and corresponding uncertainty map. The initial + threshold is ``snrthres`` :math:`=5`. + c. ``elong`` :math:`\leq` ``elongthres``: Source elongation, the ratio + A/B of semi-major to semi-minor axis. + d. (``apfluxratio`` :math:`\geq` ``apfluxratiothreslow``) and + (``apfluxratio`` :math:`\leq` ``apfluxratiothreshigh``): Ratio of flux + in a 3-pixel diameter aperture to that in a 6-pixel diameter aperture. + + .. note:: + These measurements are available in SExtractor catalogs. For + **Photutils DAOStarFinder** or another detection method, analogous + measurements will be used. DAOStarFinder's ``sharpness`` and + ``roundness`` estimates could also be used. + +2. Pixel-level metrics from a 5x5 difference-image cutout centered on each + candidate: + + a. ``nneg`` :math:`\leq` ``nnegthres``: Number of negative-valued pixels. + b. ``nbad`` :math:`\leq` ``nbadthres``: Number of pixels flagged as bad. + Currently, only pixels masked as NaN in the difference image because + of missing reference-mosaic coverage are considered bad. + c. ``sumrat`` :math:`\leq` ``sumratthres``: Apply a 3x3 median filter + with kernel truncation at cutout edges, ignoring unavailable or NaN + values. ``sumrat`` is the ratio of the sum of pixel values in the + median-filtered cutout to the sum of their absolute values. + Gaussian-distributed noise is expected to give + :math:`-0.25 \lt` ``sumrat`` :math:`\lt 0.25`, + while real signal approaches 1. -2. Pixel-level metrics based on a 5x5 cutout of the difference image centered on each candidate: - - a. ``nneg`` :math:`\leq` ``nnegthres``: Number of negative-valued pixels in the cutout. - - b. ``nbad`` :math:`\leq` ``nbadthres``: Number of pixels flagged as bad in the cutout. +3. PSF-fitting photometry quality cuts: - .. note:: - Currently "bad" pixels are only those masked as NaN in the difference image due to lack of coverage in the reference mosaic. + a. :math:`0 \lt` ``chipsf`` :math:`\lt` ``chipsfthres``: Reduced + chi-squared of the PSF fit. + b. :math:`|` ``magap3pix`` :math:`-` ``magpsf`` :math:`| \lt` + ``magdiffthres``: Difference between the appropriately + aperture-corrected 3-pixel diameter aperture magnitude and the + PSF-fit magnitude. - c. ``sumrat`` :math:`\leq` ``sumratthres``: A 3x3 median filter is applied to the cutout, with kernel trucation at the cutout edges (i.e., - unavailable or NaN pixel values are simply ignored). ``sumrat`` is defined as the ratio of the sum of pixel values in median filtered cutout - to the sum of their absolute values. For Gaussian-distributed noise, a value bewteen :math:`-0.25 \lt` ``sumrat`` :math:`\lt 0.25` is expected, - while the value approaches 1 for real signal. +4. Machine-learning real-bogus (RB) classification using all relevant + features/metrics above and available metadata, such as positional + associations with stars or galaxies in the reference image. -3. PSF-fitting photometry quality cuts: +.. _figure_of_merit: - a. :math:`0 \lt` ``chipsf`` :math:`\lt` ``chipsfthres``: The reduced chi-squared of the PSF fit. +Figure of Merit +*************** - b. :math:`|` ``magap3pix`` :math:`-` ``magpsf`` :math:`| \lt` ``magdiffthres``: Difference between the magnitude measured in a 3-pixel - diameter aperture (with appropriate aperture correction), and the PSF-fit magnitude. +The figure of merit (FOM) evaluates pipeline performance to guide algorithm +down-selection (e.g., ZOGY or SFFT for image subtraction) and parameter and +threshold tuning. It is an effective limiting magnitude for a specified +data set, combining these terms in approximate order of importance/weight: -4. Machine-learning based real-bogus (RB) classification using all relevant features/metrics described above, along with available metadata (e.g., - positional assoication with stars or galaxies in the reference image). +1. :math:`m_{\mathrm{th}}`: Average magnitude corresponding to the S/N + threshold that achieves an acceptable false-positive rate per image. +2. :math:`m_{80}`: Magnitude at which 80% of injected sources are recovered. +3. :math:`m_{20}`: Magnitude at which 20% of injected sources are recovered. +4. :math:`m_{5\sigma}`: The :math:`5\sigma` point-source limiting magnitude + on blank sky in the difference images. +5. :math:`m_{\mathrm{ph}10}`: Magnitude at which injected fluxes are + recovered with 10% precision in PSF-fitting photometry. -.. _figure_of_merit: - -Figure of Merit -********************************************** -We define a figure of merit (FOM) to evaluate the performance of the RAPID pipeline, to enable -down-selection of algorithms (e.g., ZOGY or SFFT for image subtraction), and to guide tuning of -parameters and thresholds. In broad strokes, we define the FOM as an effective limiting magnitude -for a specified data set composed of multiple terms. These are, in approximate order of general -importance/weight: (1) the average magnitude corresponding to the S/N threshold that achieves an -accepatable false-positive rate per image, :math:`m_{\mathrm{th}}`, (2) the magnitude at which 80% of -injected sources are successfully recovered, :math:`m_{80}`, (3) the magnitude at which 20% of -injected sources are successfully recovered, :math:`m_{20}`, (4) the :math:`5\sigma` point-source -limiting magnitude on blank sky, :math:`m_{5\sigma}` in the difference images, and (5) the magnitude -at which injected fluxes are recovered with 10% precision in PSF-fitting photometry, :math:`m_{\mathrm{ph}10}`. - -Each term is calculated as a weighted sum over the filters, :math:`f`, present in the test data set, -and the final FOM is then a weighted sum of each of these terms: +Each term is a weighted sum over the test data set's filters, :math:`f`; +the final FOM is a weighted sum of these terms: .. math:: \mathrm{FOM} = \left(w_{\mathrm{th}} \frac{\sum_{f} w_{\mathrm{th},f} m_{\mathrm{th},f}}{\sum_{f} w_{\mathrm{th},f}} @@ -135,90 +147,114 @@ and the final FOM is then a weighted sum of each of these terms: + w_{\mathrm{ph10}} \frac{\sum_{f} w_{\mathrm{ph10},f} m_{\mathrm{ph10},f}}{\sum_{f} w_{\mathrm{ph10},f}}\right) \\\\ / (w_{\mathrm{th}} + w_{80} + w_{20} + w_{5\sigma} + w_{\mathrm{ph10}}). -The weights, :math:`w`, can be adjusted to emphasize different aspects of the performance of the pipeline -or to prioritize performance in certain filters. Additionally, the calcuation of each term can -be further broken out by characteristics of the injected sources (e.g, transients on top of bright hosts, -hostless events, nuclear tranisents, or variables with stellar counterparts) and weighted accordingly. In -general, the RAPID pipeline needs to perform well across all Roman surveys, filters, and for a broad range -of transients and variables. +The weights, :math:`w`, can emphasize different aspects of performance or +prioritize filters. Each term can also be calculated and weighted by +injected-source characteristics, such as transients on bright hosts, +hostless events, nuclear transients, or variables with stellar +counterparts. RAPID needs to perform well across all Roman surveys and +filters, for a broad range of transients and variables. Evaluation Procedure -********************************************** -Here we describe the steps to evaluate the pipeline and calculate each term of the FOM in more detail. - -1. Define setup for evaluation and run pipeline: - - a. Define data set and injection parameters. Prior to launch, the data set could be, e.g., one test runs of the pipeline - using OpenUniverse HLTDS simulations or a RimTimSim GBTDS set (see :ref:`testing`). Post-launch, we will define specific subsets of - the survey data to use for evaluation with fake-source injection, e.g., a set fields from the HLTDS over a specific - time period. The injection parameters include the number of injections per image, their magnitude distribution, and the - specifications for the method used to assign their positions (i.e., regular grid, random positions, randomized offsets from - detected galaxies, on top of stars, etc.). - - b. Specify any pipeline modules, steps, or settings to be comparitively tested. For example, this may be the image subtraction - algorithm (ZOGY vs SFFT vs Naive), the detection method (SExtractor vs Photutils DAOStarFinder), or settings for a specific algorthm - (e.g., SEXtractor detection thresholds). - - c. Specify all axes of evalution, and their desired weights. At minimum, this includes the filters present in the data set. It may - also include properties of the injected sources (e.g., separation from host galaxy core, transients vs. variables). - - d. Run the pipeline on the test data set with the specified injections and settings to be tested. - -2. Following the completion of image subtraction and the generation of raw detection catalogs, the catalogs are filtered using the set of pre-defined - thresholds to remove the bulk of suprious candidates, as described above in . PSF-fitting photometry is then performed for all candidates - passing these criteria using the relevant difference images, corresponding uncertainty maps, and the unit-normalized PSF models computed for the - difference images. Candidates are then filtered again based on the results of PSF-fitting. - -3. Each surviving transient candidate is positionally cross-matched to the nearest source in the reference image. For simulated OpenUniverse data, - this can be done by cross-matching to galaxies brighter than ``galmatchthres`` from the simulation truth-catalogs. For real data, we will - cross-match to the reference image source catalogs generated by the pipeline, inluduing any metadata that can be used to separate stars and galaxies. - -4. Additional vetting of candidates is performed using machine-learning (ML) based real-bogus (RB) classification using all relevant features - and metadata described above. - -5. Candidates are cross-matched to the injected source catalogs (or OpenUniverse truth catalog transients) using a matching radius - of ``injmatchrad``. Candidates passing all filter criteria and matched an injected source are considered successfully recovered - true positives (TPs). Candidates passing all filter criteria, but not matched to an injected source are considered false positives (FPs). - -6. All terms of the FOM are calculated for each filter and sub-groups of injections/detetions (e.g., injections/candidates separated from the nearest - galaxy from step 3 by :math:`\gt 1.5 \times` ``injmatchrad``): - - a. :math:`m_{\mathrm{th}}`: FP candidates are grouped by :math:`m_{3\mathrm{pix}}` in :math:`\Delta m = 0.2` bins. The number - of FP candidates at or brighter than each magnitude bin, :math:`N_{\mathrm{FP}\lt m}`, is calculated. :math:`m_{\mathrm{th}}` is computed as the - magnitude where a linear interoplation of :math:`N_{\mathrm{FP}\lt m}` crosses the defined FP rate tolerance, ``fpratetol``. The S/N corresponding - to :math:`m_{\mathrm{th}}` is similarly computed via linear interpolation of the average S/N of FP candidates in each bin, and defines a threshold - ``snrfpthres`` to ensure consistency with ``fpratetol``. - - b. :math:`m_{80}` and :math:`m_{20}`: The recovery completess (fraction of injected sources in the sub-group of interest successfully recovered) - is computed in :math:`\Delta m = 0.5` bins of injected magnitude for all TP candidates passing :math:`SNR \geq` ``snrfpthres``. - The calculations accounts for the number of injected sources that lacked coverage in the reference image or fell within ``diffimedgetol`` - of any image boundary. :math:`m_{80}` and :math:`m_{20}` are computed via linear interpolation of the completness curve at 80% and 20% completeness, - respectively. - - c. :math:`m_{5\sigma}`: For each difference image, :math:`6\sigma`-clipped standard deviation of all pixel values is calculated as a - estimate of the background noise, :math:`\sigma_{\mathrm{diffbkg}}`. This value can be compared to the sigma-clipped average of the - corresponding difference uncertainty map to ensure uncertainty maps are reasonable. The :math:`5\sigma` point-source limiting flux is - then calculated as :math:`5 \sqrt{N_p} \sigma_{\mathrm{diffbkg}}`, where :math:`N_p` is the number of `noise pixels`_ of the - unit-normalized PSF model for the difference image, and converted to a limiting magnitude :math:`m_{5\sigma}` using with the appropriate - AB magnitude zeropoint. - - d. :math:`m_{\mathrm{ph}10}`: The fractional error between the recovered PSF-fit flux and the injected flux is calcualated for all TP candidates - as :math:`(f_{\mathrm{PSF}} - f_{\mathrm{inj}})/f_{\mathrm{inj}}` and grouped in :math:`\Delta m = 0.5` bins of injected magnitude. - :math:`m_{\mathrm{ph}10}` is computed via linear interpolation of the fractional errors for each bin for 10% precision. - - .. note:: - This does not account for systematic biases in the recovered fluxes, but is only an estimate of the statistical precision. - -7. The final FOM is calculated using the specified weights for each term and filter/sub-group as described above. The version of the pipeline (with choices of subtraction algorithm, - detection method, or specific parameter settings) that maximizes the FOM is judged to be better performing. +******************** + +1. Define the evaluation setup and run the pipeline: + + a. Define the data set and injection parameters. Before launch, this + could be a test run with OpenUniverse HLTDS simulations or a + RimTimSim GBTDS set (see :ref:`testing`). After launch, we will define + survey subsets for fake-source evaluation, such as HLTDS fields over + a specific period. Specify the number of injections per image, their + magnitude distribution, and the position-assignment method (regular + grid, random positions, randomized offsets from detected galaxies, + on top of stars, etc.). + b. Specify the pipeline modules, steps, or settings to compare: image + subtraction (ZOGY vs SFFT vs Naive), detection (SExtractor vs + Photutils DAOStarFinder), or algorithm settings (e.g., SEXtractor + detection thresholds). + c. Specify all evaluation axes and weights, including at least the + filters in the data set. Other axes may describe injected sources, + such as separation from the host galaxy core or transients vs. + variables. + d. Run the pipeline on the test data set with the specified injections + and settings. + +2. After image subtraction and raw-catalog generation, apply the + pre-defined filtering thresholds described above to remove most + spurious candidates. Perform PSF-fitting photometry on all survivors + using the difference images, corresponding uncertainty maps, and + unit-normalized difference-image PSF models. Filter again using the + PSF-fitting results. + +3. Positionally cross-match each surviving transient candidate to the + nearest source in the reference image. For OpenUniverse simulations, + this can use truth-catalog galaxies brighter than ``galmatchthres``. + For real data, we will use pipeline-generated reference-image source + catalogs, including metadata that can distinguish stars from galaxies. + +4. Vet candidates with machine-learning (ML) real-bogus (RB) classification + using all relevant features and metadata described above. + +5. Cross-match candidates to injected-source catalogs (or OpenUniverse + truth catalog transients) within ``injmatchrad``. Candidates passing all + filter criteria and matched to an injected source are successfully + recovered true positives (TPs); those passing all criteria but unmatched + to an injected source are false positives (FPs). + +6. Calculate all FOM terms for each filter and injection/detection + sub-group (e.g., injections/candidates separated from the nearest galaxy + from step 3 by :math:`\gt 1.5 \times` ``injmatchrad``): + + a. :math:`m_{\mathrm{th}}`: Group FP candidates by :math:`m_{3\mathrm{pix}}` + in :math:`\Delta m = 0.2` bins. Count those at or brighter than each + bin, :math:`N_{\mathrm{FP}\lt m}`. Compute :math:`m_{\mathrm{th}}` by + linearly interpolating :math:`N_{\mathrm{FP}\lt m}` to the defined FP + rate tolerance, ``fpratetol``. Interpolate the mean FP-candidate S/N + per bin to :math:`m_{\mathrm{th}}` to obtain ``snrfpthres``, the S/N + threshold consistent with ``fpratetol``. + + b. :math:`m_{80}` and :math:`m_{20}`: Calculate recovery completeness + (the fraction of injected sources recovered in the sub-group) in + :math:`\Delta m = 0.5` bins of injected magnitude, using all TPs + passing :math:`SNR \geq` ``snrfpthres``. Account for injections + lacking reference-image coverage or falling within ``diffimedgetol`` + of any image boundary. Linearly interpolate the completeness curve + to obtain :math:`m_{80}` and :math:`m_{20}` at 80% and 20% + completeness, respectively. + + c. :math:`m_{5\sigma}`: Estimate background noise, + :math:`\sigma_{\mathrm{diffbkg}}`, from the :math:`6\sigma`-clipped + standard deviation of all pixel values in each difference image. + This can be compared with the sigma-clipped average of the + corresponding difference uncertainty map to check that the map is + reasonable. Calculate the :math:`5\sigma` point-source limiting flux as + :math:`5 \sqrt{N_p} \sigma_{\mathrm{diffbkg}}`, where :math:`N_p` is + the number of `noise pixels`_ in the unit-normalized difference-image + PSF model. Convert to :math:`m_{5\sigma}` with the appropriate AB + magnitude zeropoint. + + d. :math:`m_{\mathrm{ph}10}`: For all TPs, calculate the fractional error + between recovered PSF-fit and injected flux, + :math:`(f_{\mathrm{PSF}} - f_{\mathrm{inj}})/f_{\mathrm{inj}}`, and + group it in :math:`\Delta m = 0.5` bins of injected magnitude. + Linearly interpolate the fractional errors per bin to obtain + :math:`m_{\mathrm{ph}10}` at 10% precision. + + .. note:: + This estimates statistical precision only; it does not account + for systematic biases in recovered fluxes. + +7. Calculate the final FOM with the specified weights for each term and + filter/sub-group. The pipeline version that maximizes the FOM, with its + choice of subtraction algorithm, detection method, or parameter + settings, is judged to perform better. .. _noise pixels: https://web.ipac.caltech.edu/staff/fmasci/home/mystats/noisepix_specs.pdf Evaluation of Pipeline Test runs -********************************************** +******************************** .. toctree:: :maxdepth: 2 - + openuniv_eval.rst diff --git a/docs/source/archive/archive.rst b/docs/source/archive/archive.rst index 94d38943..d46046dc 100644 --- a/docs/source/archive/archive.rst +++ b/docs/source/archive/archive.rst @@ -5,14 +5,14 @@ RAPID Archive Deliveries Introduction ************************************ -RAPID (Roman Alerts Promptly from Image Differencing) delivers prompt time-domain products and services for the Nancy Grace Roman Space Telescope. The core products are: +RAPID (Roman Alerts Promptly from Image Differencing) delivers prompt time-domain products and services for the Nancy Grace Roman Space Telescope: * **Difference images** of every new WFI SCA image against a deep reference image * **Public alert stream** of transient and variable candidates extracted from the difference images * **Light curves** (source-matched photometry) for every candidate observed more than once * **Forced photometry** at any observed sky location on request -The goal is to issue alerts within one hour of receiving L2 data from the SOC. Processing is continuous; time-critical products (alerts, difference images) are served directly by RAPID, while accumulated products are rolled up into monthly batch deliveries to the MAST archive. +RAPID processes data continuously, aiming to issue alerts within one hour of receiving L2 data from the SOC. It serves time-critical alerts and difference images directly and delivers accumulated products to the MAST archive in monthly batches. .. image:: flow.png @@ -22,7 +22,7 @@ Full pipeline documentation is available at `RAPID on ReadTheDocs ` Each reference image is accompanied by a coverage map, an uncertainty image, a PSF model, a SExtractor source catalog, and PSF-fit source and finder catalogs. -For details of the tiling scheme and reference-image construction, see :doc:`/pl/pl`. For quality analysis of current reference images, see :doc:`/prod/products`. +See :doc:`/pl/pl` for tiling and reference-image construction, and :doc:`/prod/products` for quality analysis of current reference images. Difference Images ************************************ -Each incoming SCA image is differenced against the best available reference image for its sky tile and filter. One pipeline job processes one SCA image; each 18-SCA exposure therefore generates up to 18 independent jobs. +Each incoming SCA image is differenced against the best available reference for its sky tile and filter. One job processes one SCA image, giving up to 18 independent jobs per 18-SCA exposure. The pipeline currently evaluates three differencing methods: -* **ZOGY** --- the primary method, producing a difference image, SCORR (signal-to-noise) image, uncertainty map, and PSF model. -* **SFFT** --- cross-convolution subtraction, producing decorrelated and cross-convolved images. -* **Naive** --- simple pixel-by-pixel subtraction as a diagnostic baseline. +* **ZOGY**: the primary method, producing a difference image, SCORR (signal-to-noise) image, uncertainty map, and PSF model. +* **SFFT**: cross-convolution subtraction, producing decorrelated and cross-convolved images. +* **Naive**: simple pixel-by-pixel subtraction as a diagnostic baseline. -For each method, SExtractor and PhotUtils PSF-fit catalogs are generated from both positive and negative difference images. Gain matching and sub-pixel alignment of the reference image are performed before differencing. +Gain matching and sub-pixel alignment of the reference image precede differencing. Each method generates SExtractor and PhotUtils PSF-fit catalogs from both positive and negative difference images. .. note:: - For archive delivery to MAST, only **ZOGY positive** difference-image products are planned. SFFT and Naive outputs are used internally for algorithm comparison and are excluded from the volume estimates in this document. + MAST delivery is planned for only **ZOGY positive** difference-image products. Volume estimates exclude the internal SFFT and Naive algorithm-comparison outputs. -For details on the differencing algorithms and PSF handling, see :doc:`/pl/pl`. A complete listing of all difference-image products is in :doc:`/prod/products`. +See :doc:`/pl/pl` for differencing algorithms and PSF handling, and :doc:`/prod/products` for all difference-image products. Alerts ************************************ -RAPID will follow community standards for transient alerts, packaging each event as an Apache AVRO record and publishing it via Apache Kafka. The alert stream may be split into multiple Kafka *topics* based on survey, candidate type, or other criteria. +Following community standards for transient alerts, RAPID will package each event as an Apache AVRO record and publish it via Apache Kafka. The stream may use multiple Kafka *topics* by survey, candidate type, or other criteria. -Kafka is a high-throughput messaging system optimized for streaming (hot) data. Alerts will expire from Kafka after a retention window; for long-term preservation, AVRO packets are collected into monthly tarballs and delivered to MAST. +Kafka is a high-throughput messaging system optimized for streaming (hot) data. Alerts expire after a retention window; monthly AVRO tarballs delivered to MAST provide long-term preservation. Candidate Filtering ==================== Before alert generation, spurious detections are removed through a multi-stage filtering pipeline adapted from ZTF: -1. **Catalog-level cuts** --- edge distance, signal-to-noise ratio, source elongation, aperture flux ratios -2. **Pixel-level metrics** --- negative-pixel count, bad-pixel count, median-filter sum ratio on a 5x5 cutout -3. **PSF-fit quality cuts** --- reduced chi-squared of the PSF fit, aperture-vs-PSF magnitude consistency +1. **Catalog-level cuts**: edge distance, signal-to-noise ratio, source elongation, aperture flux ratios +2. **Pixel-level metrics**: negative-pixel count, bad-pixel count, median-filter sum ratio on a 5x5 cutout +3. **PSF-fit quality cuts**: reduced chi-squared of the PSF fit, aperture-vs-PSF magnitude consistency 4. **Machine-learning real/bogus classifier** using all features above plus reference-image metadata Full details are in :doc:`/analyses/pipeline_evaluation_metrics/pipeline_evaluation_metrics`. @@ -377,7 +377,7 @@ Alert rates depend strongly on source density: * **Galactic surveys** (GBTDS, GPS): 2,000 candidates per SCA image * **Extragalactic surveys** (HLTDS, HLWAS): 500 candidates per SCA image -Over the five-year mission this yields an estimated **12.5 billion alerts** totaling **0.75 PB** at 60 KB per AVRO packet. The assumed packet size is consistent with the Rubin/LSST alert design, where uncompressed packets are up to 82 KB and gzip-compressed packets average 65 KB (`DMTN-102 `_). The final RAPID alert schema is TBD and may differ in cutout stamp and history content. The GBTDS alone accounts for 74% of all alerts, driven by the high source density in the Galactic Bulge. Alert rates are strongly seasonal, following the bulge visibility windows. +These rates yield an estimated **12.5 billion alerts** over the five-year mission, totaling **0.75 PB** at 60 KB per AVRO packet. This packet size is consistent with the Rubin/LSST alert design: up to 82 KB uncompressed and an average of 65 KB gzip-compressed (`DMTN-102 `_). The final RAPID schema is TBD; cutout stamp and history content may differ. High source density in the Galactic Bulge makes GBTDS responsible for 74% of all alerts. Rates are strongly seasonal, following bulge visibility windows. .. image:: alerts_by_ccs.png @@ -391,25 +391,20 @@ Over the five-year mission this yields an estimated **12.5 billion alerts** tota Light Curves ************************************ -RAPID builds light curves by cross-matching candidates from successive observations. The matching engine uses the Q3C spatial-indexing library in PostgreSQL with a match radius of 0.055 arcsec (half a Roman WFI pixel). Three database tables underpin the light-curve system: +RAPID builds light curves by cross-matching candidates across successive observations using PostgreSQL's Q3C spatial-indexing library. The match radius is 0.055 arcsec (half a Roman WFI pixel). Three database tables support this: -* **Sources** --- individual detections from difference-image catalogs, partitioned by processing date and SCA -* **AstroObjects** --- unique astronomical objects, partitioned by Roman-tessellation sky tile -* **Merges** --- associations linking Sources to AstroObjects +* **Sources**: individual detections from difference-image catalogs, partitioned by processing date and SCA +* **AstroObjects**: unique astronomical objects, partitioned by Roman-tessellation sky tile +* **Merges**: associations linking Sources to AstroObjects -In the 2026-07-22 large-scale test (7,272 SCA SOC-sim images), -~250 million sources were loaded into -Sources__ child database tables in 1.3 hours with 8 parallel processes -(regardless of ``flags`` value). -Source cross-matching took 35 minutes with 8 parallel processes -for ~198 million sources (with ``flags = 0``). The test covered 360 different fields. -There were ~90 million AstroObjects records and 211,394,526 Merges records loaded -into the PostgreSQL database. -Of those merges (a.k.a. lightcurve data points), 33,223 merges -resulted from cross-matching across field boundaries (i.e., the match radius can extend -across a field boundary), which is an increase of 0.0157% in terms of number of merges. +The 2026-07-22 large-scale test covered 7,272 SCA SOC-sim images across 360 different fields: -Light-curve data are stored in the PostgreSQL operations database and periodically exported to Apache Parquet for delivery to MAST. HATS partitioning for compatibility with LINCC is under evaluation. +* ~250 million sources, regardless of ``flags`` value, loaded into Sources__ child database tables in 1.3 hours with 8 parallel processes +* ~198 million sources with ``flags = 0`` cross-matched in 35 minutes with 8 parallel processes +* ~90 million AstroObjects records and 211,394,526 Merges records loaded into PostgreSQL +* 33,223 merges (a.k.a. lightcurve data points) from cross-matching across field boundaries, increasing the merge count by 0.0157%; the match radius can extend across a field boundary + +Light curves are stored in the PostgreSQL operations database and periodically exported to Apache Parquet for MAST. HATS partitioning for LINCC compatibility is under evaluation. See the Source Matching section of :doc:`/db/db` for details of the schema and partitioning strategy. @@ -430,7 +425,7 @@ Based on ZTF experience, the service will require: Delivery Schedule ************************************ -RAPID will deliver products to MAST on a **monthly** schedule throughout the nominal mission. Products are retained for a short validation window before each batch delivery. A regular monthly cadence keeps the process routine and predictable for both RAPID and MAST operations. +RAPID will deliver products to MAST **monthly** throughout the nominal mission, after a short validation window. This cadence keeps delivery routine and predictable for both RAPID and MAST operations. Each monthly delivery will include: @@ -440,7 +435,7 @@ Each monthly delivery will include: * Archived alert packets (AVRO tarballs covering the preceding month) * Forced-photometry results -Based on the provisional mission schedule, the estimated average monthly delivery is approximately **33 TB** of difference-image products, plus reference-image updates and alert-archive tarballs. The total mission archive volume is estimated to be **2.8 PB**: +The provisional schedule gives an average monthly delivery of approximately **33 TB** of difference-image products, plus reference-image updates and alert-archive tarballs. The estimated **2.8 PB** mission archive comprises: * 2.0 PB of difference-image products * 0.75 PB of archived alert packets @@ -460,11 +455,11 @@ Reprocessing RAPID plans three full reprocessings during the mission: -* **6 months** --- incorporate improved calibrations and pipeline tuning from early operations -* **2 years** --- leverage accumulated reference images and refined algorithms at mid-mission -* **End of mission (5 years)** --- produce the definitive archive with final calibrations +* **6 months**: incorporate improved calibrations and pipeline tuning from early operations +* **2 years**: use accumulated reference images and refined algorithms at mid-mission +* **End of mission (5 years)**: produce the definitive archive with final calibrations -Each reprocessing regenerates all products from the beginning of the mission. During a **6-month overlap period**, both old and new product versions coexist in MAST to allow validation and a smooth transition for archive users. After the overlap, the old versions are deleted. +Each reprocessing regenerates all products from the mission's beginning. Old and new versions coexist in MAST for a **6-month overlap period**, allowing validation and a smooth transition for archive users. Old versions are then deleted. The overlap effectively doubles the storage requirement at each reprocessing event: @@ -480,7 +475,7 @@ Products are versioned in the :doc:`operations database ` with a ``vbest Estimated AWS Costs ************************************ -The following cost estimates are based on public AWS on-demand pricing for the us-east-1 region (March 2026). All figures use list prices and simple first-order operating assumptions; they are intended as a planning guide, not as an optimized cost model. A discount variable (``AWS_DISCOUNT``) is available in ``rapid_model.py`` for modeling negotiated or reserved-instance rates. +Cost estimates use public AWS on-demand list prices for us-east-1 (March 2026) and simple first-order operating assumptions. They are a planning guide, not an optimized cost model. In ``rapid_model.py``, ``AWS_DISCOUNT`` models negotiated or reserved-instance rates. .. list-table:: AWS Cost Components :header-rows: 1 @@ -521,9 +516,9 @@ The following cost estimates are based on public AWS on-demand pricing for the u * - Average per year - $587 K -S3 storage dominates (74% of total cost) and grows steadily as the archive accumulates. Compute costs are relatively modest for routine processing but spike during the three reprocessing events, particularly the end-of-mission reprocessing which re-runs all 9.8 M SCA images. Total 5-year cost is estimated at $2.9 M at list price; the separate 5.6 PB end-of-mission overlap peak occurs after the 60 modeled mission months. Kafka/MSK is a minor fixed cost. +S3 storage accounts for 74% of total cost and grows steadily with the archive. Routine compute costs are relatively modest but spike during the three reprocessings, especially the final re-run of all 9.8 M SCA images. The $2.9 M list-price total covers 5 years; the separate 5.6 PB end-of-mission overlap peak falls after the 60 modeled mission months. Kafka/MSK is a minor fixed cost. -These estimates exclude data-transfer (egress) costs, database hosting (RDS/EC2 for PostgreSQL), and any operational overhead. Cost-reduction options such as negotiated pricing, reserved instances, Spot usage for Batch, storage-tiering choices, delivery-format tuning, and refinements to reprocessing scope have not yet been analyzed in detail and could reduce the extrapolated costs significantly. +Estimates exclude data-transfer (egress) costs, database hosting (RDS/EC2 for PostgreSQL), and operational overhead. Negotiated pricing, reserved instances, Spot usage for Batch, storage tiering, delivery-format tuning, and refined reprocessing scope could significantly reduce costs; none has been analyzed in detail. .. image:: aws_costs.png @@ -531,7 +526,7 @@ These estimates exclude data-transfer (egress) costs, database hosting (RDS/EC2 Test Data Access ************************************ -RAPID pipeline products from testing with OpenUniverse and RimTimSim simulated data are publicly available. See :doc:`/dev/tests` and :doc:`/prod/products` for full details. +Test products from OpenUniverse and RimTimSim simulated data are public; see :doc:`/dev/tests` and :doc:`/prod/products`. The latest large-scale test run (processing date 2026-02-27) is accessible at: @@ -539,4 +534,4 @@ The latest large-scale test run (processing date 2026-02-27) is accessible at: * **Logs:** ``https://rapid-pipeline-logs.s3.us-west-2.amazonaws.com/20260227/`` * **File listing:** available from the `products page `_ -Earlier test runs are available at the same bucket prefixes with their respective processing dates. See :doc:`/dev/tests` for a chronological listing of all test runs and associated pipeline improvements. +Earlier runs use the same bucket prefixes with their processing dates. See :doc:`/dev/tests` for a chronology of all runs and associated pipeline improvements.