diff --git a/docs/source/install/install.rst b/docs/source/install/install.rst index dbefcd42..b3ea6419 100644 --- a/docs/source/install/install.rst +++ b/docs/source/install/install.rst @@ -10,27 +10,26 @@ Download the source code git clone https://github.com/Caltech-IPAC/rapid -The C code in this git repo must be built, in order to run the RAPID -pipeline. Depending on whether the build is on a Mac laptop, a -Linux machine, or inside a Docker container on a Linux machine, -there are separate build scripts referred to below. +Build the C code in this git repo before running the RAPID pipeline. +Use the script below for a Mac laptop, a Linux machine, or a Docker +container on a Linux machine. -The build commands below can be repeated safely as the build scripts -remove prior build/install files before proceeding. +The build commands are safe to repeat: each script removes prior +build/install files before proceeding. -A build can take as little as 15 minutes, with most of that time spent on the GSL and +A build can take as little as 15 minutes, mostly spent on the GSL and FFTW libraries. Building C code on Mac laptop ************************************ -The script to build on a Mac laptop the C software system for the RAPID pipeline is +The Mac laptop build script is: .. code-block:: /source-code/location/rapid/c/builds/build_laptop.csh -1. Prerequisites for the build script (you may need to install brew on your Mac laptop): +1. Install the prerequisites (you may need to install brew first): .. code-block:: @@ -40,39 +39,35 @@ The script to build on a Mac laptop the C software system for the RAPID pipeline brew install libtool brew install openblas -2. Modify the following line in the build script to configure the environment within the script, - setting the absolute path of the rapid git repo: +2. Set the absolute path of the rapid git repo in the build script: .. code-block:: setenv RAPID_SW /source-code/location/rapid -3. Modify the following line in the build script to configure the PATH - environment variable within the script, ensuring that all paths to - commands like ``make``, ``gcc``, ``ls``, ``rm``, ``gfortran``, ``autoconf``, ``automake``, ``libtool``, etc. are accessible: +3. Set PATH in the build script so commands such as ``make``, ``gcc``, + ``ls``, ``rm``, ``gfortran``, ``autoconf``, ``automake`` and ``libtool`` + are accessible: .. code-block:: setenv PATH /opt/homebrew/bin:/bin:/usr/local/bin:/usr/bin:/usr/sbin:/sbin:/opt/X11/bin - You may also have to make the following symlink if the build script complains - that it cannot find libtoolize: +If the build script cannot find libtoolize, you may also need this symlink: .. code-block:: sudo ln -s /opt/homebrew/bin/glibtoolize /opt/homebrew/bin/libtoolize -4. Run the build script: +4. Run the build script. It may take minutes or hours, depending on the + Mac laptop: .. code-block:: cd /source-code/location/rapid/c/builds ./build_laptop.csh >& build_laptop.out & -The script may take some time to finish (minutes or hours depending on the Mac laptop). - -The binary executables, libraries, and include files are -installed under the following paths: +The script installs binary executables, libraries and include files under: .. code-block:: @@ -85,7 +80,7 @@ installed under the following paths: /source-code/location/rapid/c/common/fftw/include -To run a binary executable, the run-time environment must be set up with the library path, as follows: +Before running a binary executable, set the run-time library path: .. code-block:: @@ -93,42 +88,40 @@ To run a binary executable, the run-time environment must be set up with the lib .. warning:: - ``SExtractor`` is built from the source code in this build script. If it fails, - an alternate, easier method is to simply + The script builds ``SExtractor`` from source. If that build fails, + an easier alternative is: .. code-block:: brew install sex .. note:: - This build script worked successfully on a Mac laptop running macOS Monterey - with a 2.9 GHz Dual-Core Intel Core i5 processor in a previous revision where - the ``atlas`` library was required (commit 6ff4b9a2c8f796695bd9a6f7230defd85fbd32d7). - It was recently tested on a Mac laptop with M3 Max chip running macOS Sequoia 15.6.1, - and all binary executables were successfully built. (The atlas library failed to build, but the - current revision of this build script uses the ``openblas`` library instead. The build - commands for the ``atlas`` library are retained in the script because it may work on some - laptops, and it is good to keep options open.). + A previous revision requiring ``atlas`` (commit + 6ff4b9a2c8f796695bd9a6f7230defd85fbd32d7) worked on a Mac laptop + running macOS Monterey with a 2.9 GHz Dual-Core Intel Core i5 processor. + A recent test on a Mac laptop with an M3 Max chip running macOS Sequoia + 15.6.1 built all binary executables successfully. The atlas library + failed to build, but the current script uses ``openblas`` instead. + The ``atlas`` build commands remain as an option for laptops where + the library may build successfully. Building C code on Linux machine ************************************ -The script to build on a Linux machine the C software system for the RAPID pipeline is +The Linux build script is: .. code-block:: /source-code/location/rapid/c/builds/build.csh -It is assumed the atlas library is located in +The script assumes gfortran is in PATH and the atlas library is in: .. code-block:: /usr/lib64/atlas -Furthermore, it is assumed gfortran is in the PATH. - -1. Modify the following line in the build script to configure the environment within the script, setting the absolute path of the rapid git repo: +1. Set the absolute path of the rapid git repo in the build script: .. code-block:: @@ -141,8 +134,7 @@ Furthermore, it is assumed gfortran is in the PATH. cd /source-code/location/rapid/c/builds ./build.csh >& build.out & -The binary executables, libraries, and include files are -installed under the following paths: +The script installs binary executables, libraries and include files under: .. code-block:: @@ -155,41 +147,35 @@ installed under the following paths: Building C code on EC2 instance inside Docker container ************************************ -The script to build inside a Docker container the C software system for the RAPID pipeline is +The Docker container build script is: .. code-block:: /source-code/location/rapid/c/builds/build_inside_container.sh -This script has preconfigured RAPID_SW and PATH environment -variables. The former is tied directly to how the docker container is -launched, as shown in the instructions below, and the latter is tied -to how the infrastructure software in -RAPID project's Docker image has been pre-installed. +The script preconfigures RAPID_SW for the container launch shown below +and PATH for the infrastructure software pre-installed in the RAPID +project's Docker image. -1. Install ``docker`` and create Docker image if not already done +1. Install ``docker`` and create the Docker image if not already done (otherwise, skip to step 2): * How to :doc:`install Docker on EC2 instance ` * How to :doc:`create Docker image ` -2. Ssh into the EC2 instance, and launch the Docker container with the - following commands: +2. Ssh into the EC2 instance and launch the rapid:1.0 Docker image: .. code-block:: ssh -i ~/.ssh/MyKey.pem ubuntu@ubuntu@ec2-34-219-130-182.us-west-2.compute.amazonaws.com sudo docker run -it -v /source-code/location/rapid:/code rapid:1.0 bash -In this case, the rapid:1.0 Docker image is run. - -The C-code-build location is embedded in the source-code location, as -documented below. The source-code location is -mapped from a location outside the container to inside the container -in the ``docker run -v`` command option. -Therefore, the C-code build only needs to be done once, and this will -be persisted even after exiting the container. +The C-code-build location is within the source-code location, as shown +below. The ``docker run -v`` option maps that location from outside the +container to inside it, so the build needs to run only once and persists +after the container exits. The binary executables and libraries are +visible outside the container but cannot be executed there. 3. Run the build script inside the container: @@ -200,8 +186,8 @@ be persisted even after exiting the container. tail -f build_inside_container.out -The binary executables, libraries, and include files are -installed under the following paths inside the container: +The script installs binary executables, libraries and include files under +these paths inside the container: .. code-block:: @@ -213,7 +199,7 @@ installed under the following paths inside the container: /code/c/common/wcstools/wcstools-3.9.7/bin /code/c/common/wcstools/wcstools-3.9.7/libwcs -Here are listings: +Directory listings: .. code-block:: @@ -229,18 +215,15 @@ Here are listings: # ls /code/c/common/fftw/include fftw3.f fftw3.f03 fftw3.h fftw3l.f03 fftw3q.f03 -The binary executatables and libraries therein cannot be executed -outside the container even though they are visible outside. - -The wcslib library located in /code/c/lib and /code/c/include is that -of Mark M. R. Calabretta (`URL `_). +The wcslib library in /code/c/lib and /code/c/include is from +Mark M. R. Calabretta (`URL `_). -The WCS tools of Jessica Mink also has a libwcs.a (located in /code/c/common/wcstools/wcstools-3.9.7/libwcs), which may be a -different version (`URL `_). +Jessica Mink's WCS tools also provide libwcs.a, in +/code/c/common/wcstools/wcstools-3.9.7/libwcs, which may be a different +version (`URL `_). -To run a binary executable, you must first set LD_LIBRARY_PATH. Here -is an example of running ``awaicgen`` without command-line options to -get its online tutorial: +Set LD_LIBRARY_PATH before running a binary executable. This example +runs ``awaicgen`` without command-line options to get its online tutorial: .. code-block:: diff --git a/docs/source/ops/bulk_run.rst b/docs/source/ops/bulk_run.rst index ca45488f..58ae16ce 100644 --- a/docs/source/ops/bulk_run.rst +++ b/docs/source/ops/bulk_run.rst @@ -4,59 +4,54 @@ RAPID Pipeline Execution Overview ************************************ -There are many steps to running the RAPID pipeline. -It is crucial to follow the listed order below, and do not omit -any steps, as a given step relies on all previous steps. - -Here are the steps: +Run all four steps in order; each depends on the preceding steps: 1. Run science pipelines (``ppid = 15``). 2. Register pipeline metadata in operations database. 3. Run post-processing pipelines (``ppid = 17``). - Relies on updated operations database from step 2. + These rely on the operations database updated in step 2. 4. Register additional pipeline metadata in operations database. -All steps are executed within the environment of a Docker container -that is running on an EC2 instance with access to the operations database. -Any scripts to automate the execution of pipelines must -necessarily involve the ``docker run`` command. +Execute all steps in a Docker container on an EC2 instance with access to +the operations database. Automation scripts must use ``docker run``. -Launch scripts are used to run pipelines. -These query the operations database and launch pipeline-instance jobs under AWS Batch. -Individual AWS Batch jobs do not themselves interact with the operations database. +Launch scripts query the operations database and submit pipeline-instance +jobs to AWS Batch. Individual AWS Batch jobs do not interact with the +operations database. Instructions ******************************************** -The following shows commands to launch instances of the RAPID science pipeline as AWS Batch jobs -for a given range of observation dates. It is assumed that all AWS Batch jobs will finish under -the same processing date. In the example below, it is assumed the processing date is April 4, 2025 (``20250404``). - -The to-be-run-under-AWS-Batch Docker container rapid_science_pipeline:latest -self-contains a -RAPID git-clone in the /code directory, so no volume binding to an -external filesystem containing the RAPID git repo is necessary. -The container name is arbitrary, and is set to "russ-test-jobsubmit" in the example below. -Since this Docker image contains the ENTRYPOINT instruction, you must override it with the ``--entrypoint bash`` option -(and do not put ``bash`` at the end of the command). +This example launches RAPID science pipelines as AWS Batch jobs for an +observation-datetime range. It assumes processing on April 4, 2025 +(``20250404``), with all AWS Batch jobs finishing under that processing +date. Observation dates are distinct from the processing date. +The Docker container rapid_science_pipeline:latest used under AWS Batch +contains a RAPID git clone in /code; no volume binding to an external RAPID +git repository is needed. The container name is arbitrary; this example +uses "russ-test-jobsubmit". Override the image's ENTRYPOINT with +``--entrypoint bash``; do not append ``bash`` to the command. -Step 1 -============= +Log into the EC2 instance and perform Steps 1 through 4 as root +(``sudo su``). Steps 2 through 4 run inside a container with the same +environment as Step 1. -Assume we process the data on April 4, 2025 (``20250404``). This is the processing date. +After launching jobs in Steps 1 and 3, manually monitor the AWS Batch +console until all jobs complete, then proceed to registration in Steps 2 +and 4, respectively. Monitoring will be automated at a later stage of +development. -Log into EC2 instance and, from root account (``sudo su``), perform Steps 1 through 4. -Launch AWS Batch jobs for the RAPID science pipeline. +Step 1 +============= -The data to be processed are specified by the observation datetime range. -The environment variables STARTDATETIME and ENDDATETIME refer to the -start and end observation datetimes (an observation date is distinctly different from a processing date). +Launch RAPID science-pipeline jobs. STARTDATETIME and ENDDATETIME specify +the start and end observation datetimes of the data to process. .. code-block:: @@ -91,9 +86,6 @@ start and end observation datetimes (an observation date is distinctly different python3.11 /code/pipeline/awsBatchSubmitJobs_launchSciencePipelinesForDateTimeRange.py >& awsBatchSubmitJobs_launchSciencePipelinesForDateTimeRange.out & -Manually monitor the AWS Batch console to verify all jobs ran to completion. -This will be automated at some later stage of development. - RAPID products are organized in an S3 bucket according to the :doc:`processing date `. The same processing date is a required input parameter in Step 2. @@ -101,11 +93,9 @@ The same processing date is a required input parameter in Step 2. Step 2 ============ -To be executed only after all AWS Batch jobs in Step 1 have completed. - -For this step, the data to be processed are specified simply by the processing date -as an argument on the command line of the Python script ``registerCompletedJobsInDB.py``. -This is executed inside a container with the same environment as defined for Step 1. +Register science-pipeline metadata. The processing date selects the data +and is passed as a command-line argument to the Python script +``registerCompletedJobsInDB.py``. .. code-block:: @@ -117,11 +107,8 @@ This is executed inside a container with the same environment as defined for Ste Step 3 ============ -Launch AWS Batch jobs for the RAPID post-processing pipeline. - -For this step, the data to be post-processed are specified simply by the processing date -via the environment variable JOBPROCDATE (different from observation date). -This is executed inside a container with the same environment as defined for Step 1. +Launch RAPID post-processing-pipeline jobs under AWS Batch. JOBPROCDATE +selects the data by processing date, not observation date. .. code-block:: @@ -131,18 +118,13 @@ This is executed inside a container with the same environment as defined for Ste python3.11 /code/pipeline/awsBatchSubmitJobs_launchPostProcPipelinesForProcDate.py >& awsBatchSubmitJobs_launchPostProcPipelinesForProcDate_20250404.out & -Manually monitor the AWS Batch console to verify all jobs ran to completion. -This will be automated at some later stage of development. - Step 4 ============ -To be executed only after all AWS Batch jobs in Step 3 have completed. - -For this step, the data to be processed are specified simply by the processing date -as an argument on the command line of the Python script ``registerCompletedJobsInDBAfterPostProc.py``. -This is executed inside a container with the same environment as defined for Step 1. +Register post-processing metadata. The processing date selects the data +and is passed as a command-line argument to the Python script +``registerCompletedJobsInDBAfterPostProc.py``. .. code-block:: @@ -154,37 +136,40 @@ This is executed inside a container with the same environment as defined for Ste Performance ******************************************** -The AWS Batch jobs are configured to each require a machine with 4 vCPUs, 16 GB of memory, and 20 GB of disk space. -AWS Batch for the RAPID pipeline is configured to have up to 1000 jobs running in parallel, -and this can be easily increased as needed; however, the number of parallel jobs is contingent -upon the AWS Batch machine availability, which can vary with load from competing AWS customers external to the RAPID project. +Each AWS Batch job requires a machine with 4 vCPUs, 16 GB of memory, and +20 GB of disk space. RAPID's AWS Batch configuration allows up to 1000 +parallel jobs, a limit that can easily be increased. Actual concurrency +depends on machine availability, which varies with competing demand from +AWS customers outside the RAPID project. -The addition of SFFT image differencing to the science pipeline raised the machine memory requirement from 8 GB to 16 GB, -and also increased the pipeline execution time by about 3 minutes. +Adding SFFT image differencing raised the science pipeline's memory +requirement from 8 GB to 16 GB and its execution time by about 3 minutes. Step 1 ============ -On an 8-core job-launcher machine (``t3.2xlarge`` EC2 instance), it takes 1183 seconds -to launch 2069 RAPID-science-pipeline jobs with 8-core multiprocessing. +Launching 2069 RAPID-science-pipeline jobs takes 1183 seconds with 8-core +multiprocessing on an 8-core job-launcher machine (``t3.2xlarge`` EC2 +instance). -The 2069 RAPID-science-pipeline jobs take 480 seconds on average to run in parallel under AWS Batch, once -the job has actually started on the AWS Batch machine. There can, however, be significant time spent waiting -in the AWS Batch queue, as illustrated by the histogram below. Also, the pipeline itself running on an AWS Batch machine -takes longer than the reported elapsed times last month because now the pipeline computes difference images and catalogs -for both ZOGY and SFFT. -There were 80 failed pipelines because there were no prior observations for which to generate reference images. +The 2069 jobs run in parallel under AWS Batch and average 480 seconds each +once started, excluding potentially significant queue waits. Execution +takes longer than the elapsed times reported last month because the +pipeline now computes difference images and catalogs for both ZOGY and +SFFT. There were 80 failed pipelines because no prior observations were +available to generate reference images. -Here is a histogram of the AWS Batch queue wait times for an available AWS Batch machine on which to run a pipeline job: +Histogram of AWS Batch queue wait times for an available machine: .. image:: queue_wait_times.png -In theory, an AWS Batch machine with 4 vCPUs and 16 GB of memory is a scarcer resource than those with -only one vCPU and 8 GB of memory that were being used in pipeline testing last month. -That, along with potential competition from AWS customers external to the RAPID project, may explain -the relatively longer wait times in the AWS Batch queue for available machines. +In theory, machines with 4 vCPUs and 16 GB of memory are scarcer than the +machines with one vCPU and 8 GB used in pipeline testing last month. This, +along with possible competing demand from AWS customers outside RAPID, may +explain the longer queue waits. -Here is a histogram of the job execution times, measured from pipeline start to pipeline finish on an AWS Batch machine: +Histogram of job execution times, measured from pipeline start to finish +on an AWS Batch machine: .. image:: pipeline_execution_times.png @@ -192,28 +177,27 @@ Here is a histogram of the job execution times, measured from pipeline start to Step 2 ============ -On an 8-core job-launcher machine, it takes 415 seconds -to register database records for 2069 RAPID-science-pipeline jobs with 8-core multiprocessing. +Registering database records for 2069 RAPID-science-pipeline jobs takes +415 seconds with 8-core multiprocessing on an 8-core job-launcher machine. Records are inserted and/or updated in the Jobs, DiffImages, DiffImMeta, RefImages, RefImCatalogs, RefImMeta, and RefImImages database tables. -For development, the RAPID operations database is deployed on a ``t2.micro`` EC2 machine, -which has only one virtual core (1 vCPU). +The development RAPID operations database runs on a ``t2.micro`` EC2 +machine with one virtual core (1 vCPU). Step 3 ============ -On an 8-core job-launcher machine, it takes 1051 seconds -to launch 1989 RAPID-post-processing-pipeline jobs with 8-core multiprocessing. +Launching 1989 RAPID-post-processing-pipeline jobs takes 1051 seconds with +8-core multiprocessing on an 8-core job-launcher machine. The 1989 RAPID-post-processing-pipeline jobs take less than 60 seconds to run in parallel under AWS Batch. Step 4 ============ -It takes 476 seconds to register database records for 1989 RAPID-post-processing-pipeline jobs running as a single process. +Registering database records for 1989 RAPID-post-processing-pipeline jobs +takes 476 seconds as a single process. Records are updated in the Jobs, DiffImages, and RefImages database tables. - - diff --git a/docs/source/prod/products.rst b/docs/source/prod/products.rst index e92df23a..287a84b5 100644 --- a/docs/source/prod/products.rst +++ b/docs/source/prod/products.rst @@ -4,43 +4,66 @@ RAPID Pipeline Products Overview *********** -The products processed on any given date (```` Pacific Time) will be located in -the RAPID-product S3 bucket with the processing date as a prefix:: +Products are stored in the RAPID-product S3 bucket under their processing +date (```` Pacific Time):: aws s3 ls --recursive s3://rapid-product-files/ -For example, this command covers all jobs under processing date ``20260513``:: +List all jobs for processing date ``20260513``:: aws s3 ls --recursive s3://rapid-product-files/20260513 -Here are the available products for just one job (``jid=90828``) under that processing date:: +Each job differences one science image. List the products for ``jid=90828`` +on that date:: aws s3 ls --recursive s3://rapid-product-files/20260513/jid90828 -Note that there is one science image differenced per job. - -The associated product config output file is:: +The associated product config output file is parsed for metadata loaded into +the RAPID operations database after processing:: aws s3 ls --recursive s3://rapid-product-files/20260513/product_config_jid90828.ini -This is parsed for metadata to load into the RAPID operations database after the processing. +Public Access +*************** + +To download a product, construct its URL using the filename, which must be +known in advance. For example:: + + https://rapid-product-files.s3.us-west-2.amazonaws.com/20260520/jid90828/awaicgen_output_mosaic_cov_map.fits + +A full product listing is generated on demand with ``aws s3 ls``, not +committed to this repository. Redirect the command's output to recreate the +per-date listing previously distributed as a static download:: + + aws s3 ls --recursive s3://rapid-product-files/ > rapid-product-files_.txt + +A simple Python script can parse the listing to generate ``wget`` or +``curl`` download commands. + +Pipeline logs are also public, with one log file per processed science +image. The log-file URL template corresponding to the example above is:: + + https://rapid-pipeline-logs.s3.us-west-2.amazonaws.com/20260513/rapid_pipeline_job_20260513_jid90828_log.txt + -The input and intermediate files for debugging and final products are listed in the table below. -The input filenames are unique. -The product filenames are canonical and predictable: they are the same from -one science-image case to the next in different directories. +Product Files +************* + +The table lists input files, intermediate files for debugging, and final +products. Input filenames are unique. Product filenames are canonical and +predictable, repeating across science-image cases in different directories. .. warning:: - Not all products listed below may be available for a given processing date. - It depends on the particular test associated with that date. - Also, some products were newly added later. + Availability depends on the test associated with each processing date; + some products were added later. Not every date has all listed products. -There are difference-image products from three different methods: ZOGY, SFFT with cross-convolution, -and naive (simple science image minus reference image). +Difference-image products use three methods: ZOGY, SFFT with +cross-convolution, and naive (simple science image minus reference image). .. note:: - Product filenames that do not include the suffix "_negative" are for positive difference images ("science image minus reference image"). - Product filenames that include the suffix "_negative" are for negative difference images ("reference image minus science image"). + Filenames without the suffix "_negative" identify positive difference + images ("science image minus reference image"). Those with "_negative" + identify negative difference images ("reference image minus science image"). ============================================================== ======================================================================================================================= @@ -94,37 +117,12 @@ naive_masked_psfcat_residual.fits PhotUtils residu ============================================================== ======================================================================================================================= -Public Access -*************** - -To download a RAPID pipeline product, the -user must construct a URL, knowing the filename in advance, like the following:: - - https://rapid-product-files.s3.us-west-2.amazonaws.com/20260520/jid90828/awaicgen_output_mosaic_cov_map.fits - -A full listing of product files is not committed to this repository; it is -produced on demand from the bucket with the ``aws s3 ls`` command shown -above. Redirect that command's output to a file to get the same per-date -listing that was previously distributed as a static download, for example:: - - aws s3 ls --recursive s3://rapid-product-files/ > rapid-product-files_.txt - -A simple Python script can be written to parse the listing and generate ``wget`` or ``curl`` download commands. - -The pipeline log files are also publicly accessible. There is a log file for each science image processed. -Here is a template for the log-file URL that corresponds to the above example:: - - https://rapid-pipeline-logs.s3.us-west-2.amazonaws.com/20260513/rapid_pipeline_job_20260513_jid90828_log.txt - - Example Reference-Image FITS Header ****************************************** -This section lists an example reference-image FITS header to expose the user to the -various useful metadata contained therein. The keywords near the end of the listing -include operations database IDs written to the FITS header by the RAPID post-processing pipeline. - -Note that all reference images are scaled to have a fixed MAGZP of 17.0 mag. +The example header below shows reference-image metadata, including +operations database IDs written near the end by the RAPID post-processing +pipeline. All reference images are scaled to a fixed MAGZP of 17.0 mag. .. code-block:: @@ -210,8 +208,8 @@ JDEND Observation JD of latest input image used [days] MAGZP Zero point of reference image [mag] ================ ================================================================================== -Here is an image-view of the above-mentioned reference image. Note the areas of uneven coverage, -including two blue patches representing NaNs (pixels storing not a number). +The reference image above has uneven coverage, including two blue patches +representing NaNs (pixels storing not a number): .. image:: s3_rapid-product-files_20250404_jid999_awaicgen_output_mosaic_image.png @@ -219,24 +217,21 @@ including two blue patches representing NaNs (pixels storing not a number). Analysis of Reference Images ************************************ -The number of input frames that went into computing a reference image -is an important attribute of a reference image. This is listed in the -reference-image FITS header, given by FITS keyword ``NFRAMES``, along -with the filenames of the particular input images used (``INFIL###``). - -Here is a histogram of the number of input frames for our current set of 1696 reference images: +The input-frame count is an important reference-image attribute, recorded +in the FITS header as ``NFRAMES`` alongside the input filenames +(``INFIL###``). This histogram shows the counts for the current set of +1696 reference images: .. image:: rapid_refimmeta_nframes_1dhist.png -The quality-assurance metric ``cov5percent``, given by FITS keyword ``COV5PERC``, -is an absolute quantifier for the aggregate areal-depth coverage of a reference image at a -reference depth of 5, corresponding to a coadd depth of at least 5 input images. -It is computed from the reference-image coverage map. -It is defined as a percentage of the sum of the limited coverage of all pixels in an image, -where the limited coverage is all coverage and any coverage greater than 5 that is reset to 5 -for scoring purposes, relative to 5 times the total number of pixels in the image. +The quality-assurance metric ``cov5percent`` (FITS keyword ``COV5PERC``) +quantifies absolute aggregate areal-depth coverage at a reference depth of +5, corresponding to a coadd depth of at least 5 input images. Computed from +the reference-image coverage map, it is the sum of all pixel coverages, +with values greater than 5 reset to 5 for scoring, expressed as a +percentage of 5 times the total number of pixels. -Here is a histogram of cov5percent for our current set of 1696 reference images: +This histogram shows cov5percent for the current set of 1696 reference images: .. image:: rapid_refimmeta_cov5percent_1dhist.png @@ -246,119 +241,119 @@ Alerts .. warning:: - **The RAPID alert schema are under initial development** + **The RAPID alert schema is under initial development** - The records, parameter names, types, and semantics documented below are a - work in progress and may change -- including in backward-incompatible - ways -- without notice. Many parameters are currently stubs that are - always serialized as null. Do not build production consumers against - this schema yet. + Records, parameter names, types, and semantics may change without + notice, including backward-incompatible changes. Many parameters are + stubs always serialized as null. Do not build production consumers + against this schema yet. Summary ================================== -RAPID produces alerts for source detections on difference images. Each alert -is a single Apache Avro packet, assembled and serialized by the ``alerts`` -package, and consists of the following records: +Each RAPID alert reports a source detection on a difference image in a +single Apache Avro packet, assembled and serialized by the ``alerts`` +package from these records: -- ``alert`` -- the top-level record: provenance, the triggering source +- ``alert``: the top-level record: provenance, the triggering source detection, object history, and image cutouts. -- ``diaSource`` -- the triggering source detection on a difference image, +- ``diaSource``: the triggering source detection on a difference image, including astrometry, PSF-fit photometry, and fit-quality parameters. -- ``diaForcedSource`` -- forced photometry at the object position -- ``diaObject`` -- the associated astronomical object, aggregated from all +- ``diaForcedSource``: forced photometry at the object position +- ``diaObject``: the associated astronomical object, aggregated from all of its constituent detections. -- ``ssMatch`` -- an associated solar system source: will contain MPC designation, +- ``ssMatch``: an associated solar system source: will contain MPC designation, info about the position, and the predicted V-band magnitude. Current State of the Alert Schema ================================== -The schema is currently produced end-to-end for sources detected on -difference images. (see the tables in :ref:`alert-packet-contents` -for per-parameter implementation status) +Alerts are currently produced end-to-end for sources detected on difference +images. See :ref:`alert-packet-contents` for per-parameter implementation +status. ``alert`` +--------- -- The alert schema currently contains schema version information, the - triggering source detection, previous source detections, persistent - object metadata, including aggregate photometry, and cutouts at the - source position of the difference, science, and reference images. -- Cutouts are currently 129x129 pixels (~14"), but we may increase the - size if memory constraints allow. +The alert schema currently contains schema version information, the +triggering and previous source detections, persistent object metadata +including aggregate photometry, and difference, science, and reference +image cutouts at the source position. Cutouts are currently 129x129 pixels +(~14"); their size may increase if memory constraints allow. -- Before releasing version 1.0, we plan to include: +Planned before version 1.0: - - Forced photometry history (see ``diaForcedSource``) - - solar-system cross-matching - - cross-matching to the reference image SExtractor catalog - - cross-matches to other surveys (NED, Gaia). +- Forced photometry history (see ``diaForcedSource``) +- solar-system cross-matching +- cross-matching to the reference image SExtractor catalog +- cross-matches to other surveys (NED, Gaia). ``diaSource`` +------------- -- The source schema currently contain: +The source schema currently contains: - - A source ID from our pipeline - - Exposure metadata (MJD, exposure ID, SCA, exposure time, band, ...) - - The associated object ID - - Source centroid position and uncertainties - - PSF Photometry on the difference, science, and reference images - (at the difference image source centroid) - - PSF Fit quality parameters +- A source ID from our pipeline and the associated object ID +- Exposure metadata (MJD, exposure ID, SCA, exposure time, band, ...) +- Source centroid position and uncertainties +- PSF Photometry on the difference, science, and reference images at the + difference image source centroid +- PSF Fit quality parameters -- We plan to include: +Planned additions: - - Reference image ID, including co-add information - - Aperture photometry - - Shape measurements from SExtractor (currently migrating from - photutils) - - Flags, including whether the source is a likely solar-system - object +- Reference image ID, including co-add information +- Aperture photometry +- Shape measurements from SExtractor (currently migrating from photutils) +- Flags, including whether the source is a likely solar-system object ``diaForcedSource`` +------------------- -- The Forced Photometry schema is currently a stub, but will be - populated in version 1.0 -- We are benchmarking Forced Photometry routines to determine the best - algorithm. If we can successfully optimize this, we will be able to - deliver FP at alert time, but if not, data will be stored using the - first detected object position as the FP anchor. +The Forced Photometry schema is a stub to be populated in version 1.0. +Forced Photometry routines are being benchmarked to select the best +algorithm. Successful optimization would allow FP delivery at alert time; +otherwise, data will be stored using the first detected object position as +the FP anchor. -- Each Forced Source object will contain: +Each Forced Source object will contain: - - The forced photometry ID and object ID - - MJD, Exposure ID, SCA number, and band - - Measurement position - - PSF photometry on difference and science image +- The forced photometry ID and object ID +- MJD, Exposure ID, SCA number, and band +- Measurement position +- PSF photometry on difference and science image ``diaObject`` +------------- -- The object schema is half-populated so far, and is awaiting the - implementation of automatic photometry aggregation in the database. -- The object schema currently contains: +The object schema is half-populated, awaiting automatic photometry +aggregation in the database. It currently contains: - - Object ID - - Position and position uncertainty (standard deviation on detected - positions) +- Object ID +- Position and position uncertainty (standard deviation on detected + positions) -- We plan to include: +Planned additions: - - Coverage history - - Aggregate photometric statistics on each band (mean, min, max, slopes, - and number of measurements) +- Coverage history +- Aggregate photometric statistics on each band (mean, min, max, slopes, + and number of measurements) ``ssMatch`` +----------- + +This record awaits solar-system processing (KONA per-visit output). It +will contain the MPC designation, the predicted object position at the +triggering epoch, and the predicted V-band magnitude. -- Awaits solar-system processing (KONA per-visit output). Will contain - the MPC designation, the predicted position of the object at the - triggering epoch, and the predicted V-band magnitude. +Other cross-matches +------------------- -Other cross-matches: +A cross-match schema is planned for internal references and other surveys +in the top-level alert schema. The current plan is to include the top 3 +closest matches for each catalog, potentially accounting for the half-light +radius of extended sources. -- We plan to design a cross-match schema for both internal references - and other surveys, to be included in the top-level alert schema. Our - current plan is to include the top 3 closest matches for each catalog, - potentially taking into account the half-light radius for extended sources. Sample Alert Packet ================================== @@ -367,14 +362,14 @@ Sample Alert Packet **This sample file is for exploratory purposes only.** - Production-level avro files will be schema-less and will require - versioning with Confluent schema in order to use. This file's schema - may become out of date from the current alert code. + Production-level avro files will be schema-less and require versioning + with Confluent schema. This file's schema may become out of date relative + to the current alert code. :download:`sample_alert.avro ` (schema version ``00.02``). -This is a standard Avro object container file with the schema embedded, so -it can be read without any RAPID code, e.g.:: +The standard Avro object container file embeds its schema and can be read +without RAPID code:: import fastavro with open("sample_alert.avro", "rb") as f: