diff --git a/docs/source/dev/tests.rst b/docs/source/dev/tests.rst index 4c2e9841..bf24079b 100644 --- a/docs/source/dev/tests.rst +++ b/docs/source/dev/tests.rst @@ -3,19 +3,25 @@ RAPID Pipeline Testing #################################################### -Overview -************************************ +Tests are organized by processing date. Each date has a test involving +either OpenUniverse or RimTimSim simulated data. -The tests described below are organized by processing date. -On any given date will be a test involving either OpenUniverse or RimTimSim simulated data. +Tests +************************************ + +.. toctree:: + :maxdepth: 2 + openuniv_tests.rst + rimtimsim_tests.rst + socsim_tests.rst -Databases Used For Testing +Test databases ****************************************************** -One or more separate PostgreSQL databases have been set up for each different simulated dataset or -particularly distinct testing scenario: +Each simulated dataset or distinct testing scenario uses one or more +separate PostgreSQL databases: =============== =================== ===================================================================== Dataset Database name Comment @@ -30,89 +36,78 @@ RimTimSim rimtimsims3db New sims delivered on 6/22/26 (e =============== =================== ===================================================================== -Pipeline Improvements Pertinent To Testing Timeline +Pipeline changes by date ****************************************************** =============== =============================================================================================================================================================================================================================== Date Software modification =============== =============================================================================================================================================================================================================================== -5/14/2025 Added new capability to execute SFFT with the ``--crossconv`` flag -5/30/2025 Feed ZOGY computed astrometric uncertainties computed from gain-matching, instead of fixed value 0.01 pixels -5/30/2025 Shift the gain-matched reference image that is fed to ZOGY by subpixel x and y offsets computed from gain-matching -6/20/2025 Output new naive-difference-image product (``naive_diffimage_masked.fits``) -7/10/2025 Output new star finder catalog (``diffimage_masked_psfcat_finder.txt``) -7/17/2025 Input correct FWHMs when generating gain-matching SExtractor catalogs -7/17/2025 Switched around "cconv" versus "dconv" SFFT filenames in ``crossconv_flag`` logic -7/18/2025 Added new capability of fake-source injection +5/14/2025 Added support for executing SFFT with the ``--crossconv`` flag +5/30/2025 Fed ZOGY astrometric uncertainties from gain-matching instead of the fixed value 0.01 pixels +5/30/2025 Shifted the gain-matched reference image fed to ZOGY by subpixel x and y offsets from gain-matching +6/20/2025 Added naive-difference-image product (``naive_diffimage_masked.fits``) +7/10/2025 Added star finder catalog (``diffimage_masked_psfcat_finder.txt``) +7/17/2025 Used correct FWHMs when generating gain-matching SExtractor catalogs +7/17/2025 Swapped "cconv" and "dconv" SFFT filenames in ``crossconv_flag`` logic +7/18/2025 Added fake-source injection 7/23/2025 Added computation of naive-difference-image SExtractor catalog (``naive_diffimage_masked.txt``) -8/15/2025 Changes for 8/17/25 test (Big Run). Made correction to uncertainty-image formula. +8/15/2025 Corrected the uncertainty-image formula for the 8/17/25 test (Big Run). 8/15/2025 New PSF-fit catalog for SFFT difference image. 8/15/2025 Changed ``[FAKE_SOURCES] num_injections = 100, mag_min = 21.0, mag_max = 28.0``. 8/15/2025 Changed ``[PSFCAT_DIFFIMAGE] fwhm = 2.0``. -8/15/2025 Changed ``[SEXTRACTOR_DIFFIMAGE] FILTER_THRESH = 3.0, DEBLEND_NTHRESH = 32, WEIGHT_TYPE = "NONE,MAP_RMS", FILTER = "N"`` (last two parameters are overrided in code for ZOGY and SFFT SExtractor catalogs). +8/15/2025 Changed ``[SEXTRACTOR_DIFFIMAGE] FILTER_THRESH = 3.0, DEBLEND_NTHRESH = 32, WEIGHT_TYPE = "NONE,MAP_RMS", FILTER = "N"`` (last two parameters are overridden in code for ZOGY and SFFT SExtractor catalogs). 8/15/2025 Fed ZOGY ``dxrmsfin = 0.0, dyrmsfin = 0.0``. -9/8/2025 Modified to not limit the precision of (ra, dec) in PSF-fit catalogs. -9/16/2025 Added code to generate naive-difference-image PSF-fit catalogs. -9/16/2025 Added code to generate SExtractor catalogs and PSF-fit catalogs for negative difference images (ZOGY, SFFT, naive). -9/17/2025 Modified to feed ``sca_gain * exptime_sciimage`` as gain to method ``pipeline.differenceImageSubs.compute_diffimage_uncertainty``. -9/17/2025 Fixed bug: x and y subpixels offsets were swapped (adversely affected inputs to ZOGY, SFFT, and naive image-differencing). -9/25/2025 Added new method normalize_image to normalize science-image PSFs (required by ZOGY). +9/8/2025 Removed the precision limit on (ra, dec) in PSF-fit catalogs. +9/16/2025 Added generation of naive-difference-image PSF-fit catalogs. +9/16/2025 Added generation of SExtractor catalogs and PSF-fit catalogs for negative difference images (ZOGY, SFFT, naive). +9/17/2025 Fed ``sca_gain * exptime_sciimage`` as gain to method ``pipeline.differenceImageSubs.compute_diffimage_uncertainty``. +9/17/2025 Fixed swapped x and y subpixel offsets that adversely affected inputs to ZOGY, SFFT, and naive image-differencing. +9/25/2025 Added method normalize_image to normalize science-image PSFs (required by ZOGY). 10/10/2025 Added source matching within/without field boundaries to populate Sources, Merges, and AstroObjects database tables. 10/11/2025 Added methods to compute statistics for AstroObjects database tables. 10/29/2025 Set ``min_separation = 0.0`` pixels for PhotUtils catalog generation. 11/19/2025 Upgraded to SExtractor 2.28.2. -11/25/2025 Modified ``awaicgen`` for execution on Mac laptop (compiler is more strict than Linux). -12/4/2025 Explicitly cast data and uncertainty images as ndarrays when passed to PhotUtils methods (not sure whether this actually caused any problems). -12/8/2025 Fixed call to ``romanisim.psf.make_one_psf`` method after interface changed. -12/17/2025 New SFFT python module that works on rimtimsim images. +11/25/2025 Adapted ``awaicgen`` for Mac laptop execution (compiler is stricter than Linux). +12/4/2025 Explicitly cast data and uncertainty images as ndarrays when passed to PhotUtils methods (uncertain whether this had caused problems). +12/8/2025 Fixed the ``romanisim.psf.make_one_psf`` call after its interface changed. +12/17/2025 Added an SFFT python module for rimtimsim images. 12/22/2025 Adjusted ``awaicgen_num_threads = 2`` to match the number of VCPUs in the AWS Batch machines used by the RAPID pipeline. -1/14/2026 Modified science pipeline to output catalogs in parquet format. +1/14/2026 Changed the science pipeline to output catalogs in parquet format. 1/24/2026 Added methods to delete not-best records in Sources and Merges database tables. -1/30/2026 Developed code to generate sources and lightcurves HATS catalogs. -1/31/2026 Various miscellaneous improvements such as modifications to run RAPID science pipeline on Mac laptop. +1/30/2026 Added generation of sources and lightcurves HATS catalogs. +1/31/2026 Made miscellaneous improvements, including RAPID science-pipeline execution on Mac laptop. 2/3/2026 Created forced-photometry backend and added ``cforcepsfaper`` C module. 2/4/2026 Reduced-chi2 in PhotUtils catalogs and Sources database table. 2/11/2026 Scaled reference-image inputs so that reference image has fixed zero point = 17 mag. -2/12/2026 Modified to generate PhotUtils catalog for reference image. -3/24/2026 Completed upgrade/debugging fake-source injection for variable sources with fixed sky positions. -4/7/2026 Modified SFFT code to output a difference-image PSF. -4/9/2026 Changes to how the uncertainty images are calculated (for science image and refimage inputs). -4/13/2026 Replaces hard-wired value 1750.0 with ``saturation_value_rate_sciimage`` for processing rimtimsims. -4/16/2026 Modified crossMatchSources.py to only cross-match sources with ``flags = 0``. -4/17/2026 Modified crossMatchSources.py to cross-match using AstroObjects ``(meanra,meandec)`` instead of ``(ra0,dec0)``. -4/17/2026 Modified crossMatchSources.py to update AstroObjects ``(meanra,meandec)`` record for each lightcurve data point added. -4/20/2026 Modified to cross-match all sources in one observation at a time for all SCAs in ascending time order. -4/20/2026 Modified to load into RAPID operations database the SFFT-difference-image PhotUtils catalogs, instead of ZOGY. -4/21/2026 Modified to replace NaNs, if any, in SFFT difference image with zeros. -4/21/2026 Modified to replace NaNs, if any, in difference-image uncertainty images with ``std_dif_img``. +2/12/2026 Added PhotUtils catalog generation for the reference image. +3/24/2026 Completed upgrading and debugging fake-source injection for variable sources with fixed sky positions. +4/7/2026 Changed SFFT code to output a difference-image PSF. +4/9/2026 Changed uncertainty-image calculations (for science image and refimage inputs). +4/13/2026 Replaced hard-wired value 1750.0 with ``saturation_value_rate_sciimage`` for processing rimtimsims. +4/16/2026 Changed crossMatchSources.py to cross-match only sources with ``flags = 0``. +4/17/2026 Changed crossMatchSources.py to cross-match using AstroObjects ``(meanra,meandec)`` instead of ``(ra0,dec0)``. +4/17/2026 Changed crossMatchSources.py to update AstroObjects ``(meanra,meandec)`` record for each lightcurve data point added. +4/20/2026 Changed cross-matching to process all sources in one observation at a time for all SCAs in ascending time order. +4/20/2026 Changed RAPID operations database loading to use SFFT-difference-image PhotUtils catalogs instead of ZOGY. +4/21/2026 Replaced NaNs, if any, in SFFT difference image with zeros. +4/21/2026 Replaced NaNs, if any, in difference-image uncertainty images with ``std_dif_img``. 4/21/2026 Increased ``[SCI_IMAGE] saturation_level`` from 100000 to 1100000 for rimtimsims. -4/22/2026 Modified SFFT command for rimtimsims to use the brute-force masking options (``--bsmaskvalue 20000.0 --bsmaskradius 30.0``). -4/22/2026 In the latest version of PhotUtils, output column name ``npixfit`` has been changed to ``n_pixels_fit``, and output column name ``npix`` has been changed to ``n_pixels``. -4/23/2026 Modified to use ``filename_sfftdiffpsf`` for SFFT-difference-image PSF-fit catalog generation, instead of ``filename_refimage_psf`` as before. +4/22/2026 Changed the SFFT command for rimtimsims to use the brute-force masking options (``--bsmaskvalue 20000.0 --bsmaskradius 30.0``). +4/22/2026 The latest PhotUtils version renamed output columns ``npixfit`` to ``n_pixels_fit`` and ``npix`` to ``n_pixels``. +4/23/2026 Used ``filename_sfftdiffpsf`` for SFFT-difference-image PSF-fit catalog generation, instead of ``filename_refimage_psf`` as before. 4/24/2026 Changed ``[SOURCE_MATCHING] match_radius`` to 0.00001528 degrees (half a Roman WFI pixel). Reran cross-matching for the 4/23/2026 test. -4/28/2026 Modified SFFT code to refactor bright star masking in ``bkg_mask`` to use binary_dilation. -4/28/2026 Modified SFFT code to replace the slow per-pixel distance loop with ``scipy.ndimage.binary_dilation`` and a precomputed circular footprint. -4/28/2026 Modified SFFT code so that when a SExtractor catalog is provided, a second pass after catalog masking to catch any remaining bright pixels above ``bsmask_value``. -4/29/2026 Modified SFFT code to fix logic path issues, and set ``sat_value`` and ``bsmask_value`` defaults to 1e6 to disable masking unless explicitly set. -5/12/2026 Modified to scale the reference-image uncertainty map by the gain-matching scale factor (prior to this, gain-matching was only applied to the reference image). -5/12/2026 Moved the block of code that uploads intermediate products to just before ZOGY execution (this facilitates running ZOGY offline from S3-bucket downloaded inputs). -5/19/2026 Modified to feed ZOGY scaled std_ref_img by scalefacref (gain-matching correction). -5/27/2026 Modified gain-matching to fall back on scalefac = 10**(0.4*(magzpref - magzpsci)) if unable to find a sufficient number of matches. +4/28/2026 Refactored SFFT bright star masking in ``bkg_mask`` to use binary_dilation. +4/28/2026 Replaced SFFT's slow per-pixel distance loop with ``scipy.ndimage.binary_dilation`` and a precomputed circular footprint. +4/28/2026 Added an SFFT second pass after catalog masking, when a SExtractor catalog is provided, to catch remaining bright pixels above ``bsmask_value``. +4/29/2026 Fixed SFFT logic path issues and set ``sat_value`` and ``bsmask_value`` defaults to 1e6 to disable masking unless explicitly set. +5/12/2026 Scaled the reference-image uncertainty map by the gain-matching scale factor (prior to this, gain-matching was only applied to the reference image). +5/12/2026 Moved intermediate-product uploads to just before ZOGY execution to facilitate offline ZOGY runs with S3-bucket downloaded inputs. +5/19/2026 Fed ZOGY scaled std_ref_img by scalefacref (gain-matching correction). +5/27/2026 Changed gain-matching to fall back on scalefac = 10**(0.4*(magzpref - magzpsci)) if unable to find a sufficient number of matches. 6/2/2026 Updated L2Files database table for 5th-order SIP distortion. 6/23/2026 Upgraded awaicgen version from 5.2 to 5.4 (fixed bug in handling SIP distortion). -7/8/2026 Modified to replace NaNs in difference-image inputs with zeros. -8/5/2026 Made correction to ``data_unc`` formula (line 64 of ``pipeline/differenceImageSubs.py``). -8/5/2026 Modified to feed ``dfis.compute_diffimage_uncertainty method`` background-subtracted ``filename_bkg_subbed_science_image`` instead of ``reformatted_science_image_filename`` (with background). +7/8/2026 Replaced NaNs in difference-image inputs with zeros. +8/5/2026 Corrected the ``data_unc`` formula (line 64 of ``pipeline/differenceImageSubs.py``). +8/5/2026 Fed ``dfis.compute_diffimage_uncertainty method`` background-subtracted ``filename_bkg_subbed_science_image`` instead of ``reformatted_science_image_filename`` (with background). 8/5/2026 Set ``min_separation = 1.0`` pixels for PhotUtils catalog generation. =============== =============================================================================================================================================================================================================================== - - -Tests -************************************ - -.. toctree:: - :maxdepth: 2 - - openuniv_tests.rst - rimtimsim_tests.rst - socsim_tests.rst diff --git a/docs/source/sysarch/comp_arch.rst b/docs/source/sysarch/comp_arch.rst index effa66a4..399e3139 100644 --- a/docs/source/sysarch/comp_arch.rst +++ b/docs/source/sysarch/comp_arch.rst @@ -5,74 +5,66 @@ RAPID Computing Architecture System Architecture ************************** -Here is a high-level flowchart of the RAPID system architecture: +RAPID runs entirely in the AWS cloud and is accessible from a laptop with +an Internet connection. The high-level system architecture is shown below: .. image:: sysarch.png -Everything is done in the AWS cloud, and is accessible via a laptop with Internet connection. - -The database server is a very inexpensive ``t2.micro`` AWS machine, -which runs 24 hours a day, seven days a week. -A more powerful multi-core, high-memory machine can be utilized to execute RAPID pipeline instances, -which would only be activated as needed in order to save money. This is just one way the -RAPID pipelines could be run in parallel; the next section explains another way, which has -been demonstrated to be a very feasible methodology indeed. - -Our strategy for software interactions with the SQL database is that queries are executed only during -initial pipeline launching and final data aggregation stages, before and after pipeline instances are -executed separately on multiple CPU cores. This ensures scalability -of the RAPID-pipeline computing system. +The database server is a very inexpensive ``t2.micro`` AWS machine running +24 hours a day, seven days a week. Pipeline instances could run in parallel +on a more powerful multi-core, high-memory machine, activated only as needed +to save money. The next section describes the tested alternative: parallel +processing across multiple machines with the AWS Batch Service. Computing Architecture ************************** -Here is a another flowchart of the RAPID computing architecture, which details how -parallel processing of the data is done on multiple machines, which is has proven -to be a very viable and practical approach from our extensive testing thus far: +The AWS Batch Service enables massively parallel data processing across +multiple machines. Extensive testing has demonstrated that this approach +is viable and practical: .. image:: computing_architecture.png -Parallel processing on a massive scale is facilitated by the AWS Batch Service. - -Database interactions with a PostgreSQL database are done only during initial pipeline launching -and final data aggregation stages, before and after pipeline instances are executed under the -AWS Batch Service. This ensures scalability of the RAPID-pipeline computing system. +To ensure scalability, SQL queries and other interactions with the PostgreSQL +database occur only during initial pipeline launching and final data +aggregation, before and after pipeline instances execute separately on +multiple CPU cores or under the AWS Batch Service. Pipeline Performance ************************** -In one of our initial large-scale tests, -RAPID pipeline instances were launched for all OpenUniverse simulated images with ``DATE-OBS >= 2028-09-07 00:00:00`` -and ``DATE-OBS <= 2028-09-08 08:30:00``. This is about 2000 jobs, one job per science image. All jobs were successfully run, -except for 80 jobs in which a reference image could not be made due to lack of prior observations for the associated field. -The elapsed execution time for a RAPID pipeline job was measured -from the time it was launched to the time it finished running on an AWS Batch machine, of course, after writing -the pipeline products to the output S3 bucket. - - .. warning:: The performance results below are obsolete, but kept for historical reasons. The latest performance result can be found :doc:`here `. +An initial large-scale test launched RAPID pipeline instances for all +OpenUniverse simulated images with ``DATE-OBS >= 2028-09-07 00:00:00`` +and ``DATE-OBS <= 2028-09-08 08:30:00``: about 2000 jobs, one per science +image. All succeeded except 80 jobs that could not generate a reference +image because the associated field lacked prior observations. -Here is a histogram of the job execution times: +Elapsed execution time was measured from job launch to completion on an +AWS Batch machine, including writing the pipeline products to the output +S3 bucket. The histogram below shows these times: .. image:: rapid_job_elapsed_vs_time_1dhist.png -Here is a 2-D histogram of the job execution times versus number of input frames for the reference image that was generated: +The 2-D histogram shows job execution time versus the number of input +frames used to generate the reference image: .. image:: rapid_job_elapsed_vs_nframes_2dhist.png -It is obvious from the figure that the execution times have a contribution that is proportional -to the number of reference-image inputs. +The figure shows an execution-time contribution proportional to the number +of reference-image inputs. -The products from this test run are in the following S3 bucket:: +The test products are in the following S3 bucket:: aws s3 ls --recursive s3://rapid-product-files/20250304 -For example, here is the S3 bucket path to a SourceExtractor catalog made from the difference image for one job (jid=999):: +For example, the SourceExtractor catalog made from the difference image +for job jid=999 is at:: s3://rapid-product-files/20250304/jid999/diffimage_masked.txt diff --git a/docs/source/todo/todo_list.rst b/docs/source/todo/todo_list.rst index bbe3e178..5fd2ccfa 100644 --- a/docs/source/todo/todo_list.rst +++ b/docs/source/todo/todo_list.rst @@ -1,66 +1,77 @@ To-Do List #################################################### - Pipeline Software ************************************* -+-----------------+----------------------------------------------------------+-------+ -|Task | Comment | Done? | -+=================+==========================================================+=======+ -| Ingest L2 files | L2 files from the SOC will be in ASDF format. We need | No | -| | to copy them to our S3 bucket and register records in | | -| | the L2 files table of our operations database. | | -+-----------------+----------------------------------------------------------+-------+ -| QA system | Continue improving and refining. | Basic | -+-----------------+----------------------------------------------------------+-------+ -| Parallel DB | Modify registerCompletedJobsInDB.py to | Yes | -| record insert | register database records in parallel. | | -+-----------------+----------------------------------------------------------+-------+ -| Solar system | Joe Masiero's code needs to be integrated | No | -| Objects | into the science pipeline. | | -+-----------------+----------------------------------------------------------+-------+ -| SFFT crossconv | Propagate from configuration file. | Yes | -| flag | | | -+-----------------+----------------------------------------------------------+-------+ -| FILEDATE | Update DATE value in header of all FITS-file products. | Yes | -| needed | | | -+-----------------+----------------------------------------------------------+-------+ -| Position | Compute median offset in delta x and y from sci/ref | Yes | -| refinement | cross-matched isolated sources. Use to do orthogonal | | -| for ZOGY | subpixel-shifting of reference image data prior to ZOGY. | | -+-----------------+----------------------------------------------------------+-------+ -| Astrometric | Compute RMS delta x and y from sci/ref | Yes | -| error | cross-matched isolated sources to feed as ZOGY input | | -| for ZOGY | | | -+-----------------+----------------------------------------------------------+-------+ -| Gain- | Compute image-data gain-match scale factor from sci/ref | Yes | -| matching | cross-matched isolated sources to scale reference image | | -| for ZOGY | data prior to ZOGY. | | -+-----------------+----------------------------------------------------------+-------+ +.. list-table:: + :header-rows: 1 + :widths: 22 68 10 + * - Task + - Comment + - Done? + * - Ingest L2 files + - L2 files from the SOC will be in ASDF format. Copy them to our S3 + bucket and register records in the L2 files table of our operations + database. + - No + * - QA system + - Continue improving and refining. + - Basic + * - Parallel DB record insert + - Modify registerCompletedJobsInDB.py to register database records + in parallel. + - Yes + * - Solar system Objects + - Integrate Joe Masiero's code into the science pipeline. + - No + * - SFFT crossconv flag + - Propagate from the configuration file. + - Yes + * - FILEDATE needed + - Update the DATE value in the header of all FITS-file products. + - Yes + * - Position refinement for ZOGY + - Compute median offsets in delta x and y from sci/ref cross-matched + isolated sources. Use them for orthogonal subpixel-shifting of + reference image data before ZOGY. + - Yes + * - Astrometric error for ZOGY + - Compute RMS delta x and y from sci/ref cross-matched isolated + sources as ZOGY input. + - Yes + * - Gain-matching for ZOGY + - Compute the image-data gain-match scale factor from sci/ref + cross-matched isolated sources to scale reference image data + before ZOGY. + - Yes Operations Database ************************************* -+-----------------+----------------------------------------------------------+-------+ -|Task | Comment | Done? | -+=================+==========================================================+=======+ -| Database | Need automated software in place to periodically | No | -| backups | backup the operations database | | -| | | | -+-----------------+----------------------------------------------------------+-------+ -| Database | Need to select a powerful enough EC2 machine that runs | No | -| for pipeline | 24/7 (consider cost), and set up PostgreSQL database on | | -| operations | it for pipeline operations. Must have multiple cores. | | -| | Cost it out on AWS. | | -+-----------------+----------------------------------------------------------+-------+ -| New Fields | Nice to have would be a Fields table indexed by field | No | -| table | that has (ra,dec) of center and 4 corners of tile. | | -| | Get data from our SQLite Roman tesselation database. | | -+-----------------+----------------------------------------------------------+-------+ -| New columns in | dxrmsfin,dyrmsfin,dxmedianfin,dymedianfin | Yes | -| DiffImMeta | | | -| table | | | -+-----------------+----------------------------------------------------------+-------+ +.. list-table:: + :header-rows: 1 + :widths: 22 68 10 + + * - Task + - Comment + - Done? + * - Database backups + - Put automated software in place to periodically back up the + operations database. + - No + * - Database for pipeline operations + - Select a sufficiently powerful, multicore EC2 machine to run 24/7, + cost it out on AWS, and set up a PostgreSQL database on it for + pipeline operations. + - No + * - New Fields table + - Nice to have: a Fields table indexed by field, with (ra,dec) of + the tile center and 4 corners. Get the data from our SQLite Roman + tesselation database. + - No + * - New columns in DiffImMeta table + - dxrmsfin,dyrmsfin,dxmedianfin,dymedianfin + - Yes