diff --git a/docs/source/db/db.rst b/docs/source/db/db.rst index 7216ca8d..b81ee965 100644 --- a/docs/source/db/db.rst +++ b/docs/source/db/db.rst @@ -4,15 +4,13 @@ RAPID Operations Database Introduction ************************************ -The RAPID pipeline utilizes a PostgreSQL database. The Q3C library -has been installed as a plug-in for fast queries on sky position. - -.. note:: - The database design described below is evolving and subject to change. +The RAPID pipeline uses a PostgreSQL database with the Q3C library +installed as a plug-in for fast sky-position queries. .. note:: This page describes the database design as built on the ``dev`` - branch. On the ``rebuild`` branch, the schema is versioned SQL under + branch and is evolving and subject to change. On the ``rebuild`` + branch, the schema is versioned SQL under ``database/migrations/`` in the RAPID git repository, applied in filename order by ``database/apply-migrations.sh``; see ``database/README.md`` for the rules and how to run the applier @@ -22,51 +20,54 @@ has been installed as a plug-in for fast queries on sky position. Schema ************************************ -A diagram of the database-table schema is given as follows: +The database-table schema is shown below: .. image:: dbschema.png -There are multiple provisions for indexing on sky position: +Sky positions are indexed in several ways: * Q3C indexing -* The field column in various tables stores the Roman tessellation index for the sky tile associated with sky position. - For the RAPID project, the Roman-tessellation parameter setting NSIDE=512 will be used, - which results in tile sizes somewhat smaller than that of a Roman SCA image, - and a total of 6,291,458 tiles covering the entire sky. -* Healpix level-6 index (hp6), with an approximate resolution of 0.92 degrees (almost the width of the Roman WFI or 6 SCAs plus gaps). - There are 49,152 level-6 indices. -* Healpix level-9 index (hp9), with an approximate resolution of 0.11 degrees (almost the width of a Roman SCA). - There are 3,145,728 level-9 indices. +* The field column in various tables stores the Roman tessellation index + of the sky tile containing the position. RAPID will use the + Roman-tessellation parameter NSIDE=512, giving 6,291,458 tiles across + the sky, each somewhat smaller than a Roman SCA image. +* Healpix level-6 index (hp6): 49,152 indices with an approximate + resolution of 0.92 degrees, almost the width of the Roman WFI + (6 SCAs plus gaps). +* Healpix level-9 index (hp9): 3,145,728 indices with an approximate + resolution of 0.11 degrees, almost the width of a Roman SCA. -The L2Files database table has the ``overlapfields`` int[] column for a storing list of field numbers that a given -Roman SCA image, with its unique orientation on the sky, for fields that it overlaps. The algorithm that computes -the overlapping fields omits fields with less than 25 pixels of overlap. +The L2Files ``overlapfields`` int[] column lists the field numbers +overlapped by a Roman SCA image at its particular sky orientation. +The overlap algorithm omits fields with less than 25 pixels of overlap. Record Versioning ************************************ -L2 files, difference images, and reference images are versioned in their -respective database tables (L2Files, DiffImages, and RefImages), given by the version column. The version -is also embedded in the filesystem paths of the corresponding data files. -The best version is given by vbest, a smallint table -column that stores 0 for not best, 1 for best that is usually the -latest version, or 2 if the version is locked. It is a matter of -policy whether old versions will be kept in the filesystem and/or -database (these could be removed at will). +L2 files, difference images, and reference images carry a version in +their respective tables (L2Files, DiffImages, and RefImages), in the +version column and in the corresponding data files' filesystem paths. +The smallint column vbest identifies the best version: + +* 0: not best +* 1: best, usually the latest version +* 2: locked version + +Policy determines whether old versions remain in the filesystem and/or +database; they can be removed at will. Sky-Position Queries Using Q3C Library Functions -************************************ +********************************************** -The L2FileMeta and DiffImages database tables store the image centers -(ra0, dec0) and their four corners (rai, deci, i=1,...,4). -Database queries involving Q3C functions like the following can find all images that -overlap a given image and acquired before the image of interest, -such as the one with rid = 152336 (rid = L2File primary key), where the (ra, dec) values -below are for that image's center and four corners: +L2FileMeta and DiffImages store image centers (ra0, dec0) and four corners +(rai, deci, i=1,...,4). Q3C queries can find all images that overlap a +given image and were acquired before it. This example uses rid = 152336 +(rid = L2File primary key); the (ra, dec) values are that image's center +and four corners: .. code-block:: @@ -85,7 +86,7 @@ below are for that image's center and four corners: order by dist; -Once the relevant rids are found, the filenames can be looked up as follows: +Use the relevant rids to look up filenames: .. code-block:: @@ -98,7 +99,7 @@ Once the relevant rids are found, the filenames can be looked up as follows: Reference-Image QA ************************************ -The RefImMeta database table stores various QA measures for reference images. +RefImMeta stores reference-image QA measures: +--------------------+-----------------------------------------------------------------------------------+ | Database column | Definition | @@ -142,19 +143,18 @@ The RefImMeta database table stores various QA measures for reference images. | nsexcatsources | Number of sources in RefImage SourceExtractor catalog | +--------------------+-----------------------------------------------------------------------------------+ -The quality-assurance metric ``cov5percent``, given by FITS keyword ``COV5PERC``, -is an absolute quantifier for the aggregate areal-depth coverage of a reference image at a -reference depth of 5, corresponding to a coadd depth of at least 5 input images. -It is computed from the reference-image coverage map. -It is defined as a percentage of the sum of the limited coverage of all pixels in an image, -where the limited coverage is all coverage and any coverage greater than 5 that is reset to 5 -for scoring purposes, relative to 5 times the total number of pixels in the image. +The quality-assurance metric ``cov5percent`` (FITS keyword ``COV5PERC``) +is an absolute measure of aggregate areal-depth coverage at a reference +depth of 5, corresponding to a coadd depth of at least 5 input images. +It is computed from the reference-image coverage map: cap each pixel's +coverage at 5, sum the capped values, and express the result as a +percentage of 5 times the total number of image pixels. Difference-Image QA ************************************ -The DiffImMeta database table stores various QA measures for difference images. +DiffImMeta stores difference-image QA measures: +--------------------+-------------------------------------------------------------------------------------------+ | Database column | Definition | @@ -179,102 +179,105 @@ The DiffImMeta database table stores various QA measures for difference images. Source Matching ************************************ -Four basic PostgreSQL database tables are used for source cross-matching PSF-fit catalogs -made by the Python photutils package from the SFFT difference images and curating -source-extracted lightcurves (until a final decision on which image-differencing and -source-extraction methods are best): +Four PostgreSQL tables support cross-matching sources from PSF-fit +catalogs made by the Python photutils package from SFFT difference +images, and curating source-extracted lightcurves. These methods are used +until a final decision on the best image-differencing and source-extraction +methods: * Sources (extracted/selected from catalogs) * AstroObjects (astronomical objects for which time-dependent sources form light curves) * Merges (associations between Sources and AstroObjects via source cross-matching) * AstroObjectsMeta (statistics on astronomical-object lightcurves added after source matching) -A diagram of the source-matching database-table schema is given as follows: +The source-matching schema is shown below: .. image:: source_matching.png -As indicated in the diagram, there will be several Sources tables -named differently, according to the observing-date and SCA parameters. -Same for Merges, AstroObjects, and AstroObjectsMeta tables, according to field number. -This is to partition the data into manageable chunks. -The partitioning schemes for these tables are discussed below in more detail. - -The parent or prototype tables have the generic names: Sources, Merges, AstroObjects, and AstroObjectsMeta. -No actual records are stored in prototype tables. - -Database-table inheritance is or can be used to tie child tables, which store the actual records, -to the parent table. -At this time, only the Sources tables utilize inheritance. This is because the source ID -in the Merges table can be most easily associated with a record in the correct child-table name -by querying the Sources parent table. - -A Sources child table is created for each observation date and SCA. -Thus the partitioning scheme for sources is by time and chip number. -This design strikes a balance between partitioning for parallel processing -and non-proliferation of Sources tables. It is anticipated that most of -sources to be matched are spurious, and so this partitioning scheme will -have to be reassessed after the actual number of sources involved can be -better estimated. -PhotUtils-catalog source extractions are loaded into the Sources tables via -parallel processes in observation-date-time order. -This includes all sources, regardless of their bit-wise ``flags`` attribute. - -AstroObjects and AstroObjectsMeta database tables are created for each Roman-tessellation sky tile or field. -Merges tables are also created for each Roman-tessellation sky tile or field. -Thus the partitioning scheme for astronomical objects and associated cross-matching with -sources (via Merges tables) are by sky position. - -A unique index, called ``aid``, for each AstroObjects_ database record is computed, -not via a database sequence, but by a deterministic method that scales (ra, dec) to have -1/3300-arcsecond precision and then concatenates these scaled sky coordinates together. -This index fits within an ``int64`` data type. - -Sources and AstroObjects database tables are cross-matched for the appropriate partitions, -in observing-time order, using the join function from the Q3C-library PostgreSQL extension, -and records in the associated Merges tables are then populated. -Only sources with ``flags = 0`` are considered. -A given Sources child table can contain records for different fields, filters, and exposures. -The cross-matching is done for all sources in one observation at a time, for all SCAs, in ascending time order. - -.. note:: - Sources that are NOT matched become new records in the AstroObjects tables. - -Source matching is done in parallel by field. Thus multiple cores -on the database-server machine will be utilized, and scaling up the architecture is possible -by moving the database server to a machine with more cores and memory (as can be afforded). -Cross-matching the sources results in records loaded into the Merges_ and -AstroObjects_ database tables. -The source cross-matching extends across field boundaries for sources near field edges. - -The RAPID pipeline makes PSF-fit catalogs for both positive difference images (i,e, "science image -minus reference image") and negative difference image (i,e, "reference image minus science image"). -The Sources database table has the boolean ``isdiffpos`` column to indicate for a given source -from which type of source extraction it originated. - -For the 7/22/2026 test with SOC sims, ~250 million sources were loaded into -Sources__ child database tables in 1.3 hours with 8 parallel processes -(regardless of ``flags`` value). -Source cross-matching took 35 minutes with 8 parallel processes -for ~198 million sources (with ``flags = 0``). The test covered 360 different fields. -A match radius of 0.055 arcsec or half a Roman WFI pixel was used. -There were ~90 million AstroObjects records and 211,394,526 Merges records loaded -into the PostgreSQL database. Of those merges (a.k.a. lightcurve data points), 33,223 merges -resulted from cross-matching across field boundaries (i.e., the match radius can extend -across a field boundary), which is an increase of 0.0157% in terms of number of merges. - -The lightcurve statistics are stored in the AstroObjectsMeta_ database tables, and are inserted after the -cross-matching. This is done as a separate process, after the source cross-matching. -The script that computes the lightcurve statistics drops all -AstroObjectsMeta_ database tables and then recreates them, before the statistical computations. -The AstroObjectsMeta_ database tables are explicitly vacuumed and analyzed at the end of this process. -For the 7/22/2026 test with SOC sims, it took 1.6 hours with 8 parallel processes to compute -statistics for ~90 million AstroObjects. - -Because reprocessing generates new product versions (usually latest is best), there are -separate processes that remove not-best lightcurve data points from the Sources and Merges_ database tables, -and then explicitly clusters, vacuums, and analyzes these database tables. - -The Sources database table has the ``rb`` float column for storing real-bogus scores, -computed from a Machine-Learning algorithm that is optimized for Roman WFI data, -which is a fractional number in the [0.0, 1.0] range that corresponds to the likelihood -that a source is real (as opposed to bogus). +Partitioning +============ + +The parent or prototype tables, Sources, Merges, AstroObjects, and +AstroObjectsMeta, contain no actual records. Records reside in child +tables, partitioned into manageable chunks: + +* Sources child tables are created and named by observation date and SCA + (time and chip number). Each can contain different fields, filters, + and exposures. This balances parallel processing against table + proliferation. Most sources to be matched are expected to be spurious; + the scheme will need reassessment when source counts are better estimated. +* Merges, AstroObjects, and AstroObjectsMeta tables are created for each + Roman-tessellation sky tile or field and named by field number. Objects + and their source associations are therefore partitioned by sky position. + +Inheritance can tie child tables to their parents; currently only +Sources uses it. Querying the Sources parent is the easiest way to +associate a source ID in Merges with a record in the correct child table. + +Each AstroObjects_ record has a unique index, ``aid``, computed +deterministically rather than through a database sequence. The method +scales (ra, dec) to 1/3300-arcsecond precision and concatenates the scaled +coordinates. The result fits in an ``int64`` data type. + +Source Loading and Attributes +============================= + +PhotUtils-catalog source extractions are loaded into Sources in parallel, +in observation-date-time order, regardless of their bit-wise ``flags`` +attribute. + +RAPID makes PSF-fit catalogs for both positive difference images +("science image minus reference image") and negative difference images +("reference image minus science image"). The Sources boolean column +``isdiffpos`` records which extraction type produced each source. + +The Sources float column ``rb`` stores real-bogus scores from a +Machine-Learning algorithm optimized for Roman WFI data. A score is a +fractional number in [0.0, 1.0] corresponding to the likelihood that the +source is real rather than bogus. + +Cross-Matching +============== + +The Q3C-library PostgreSQL extension's join function cross-matches Sources +and AstroObjects in the appropriate partitions. Matching considers only +sources with ``flags = 0`` and proceeds one observation at a time, across +all SCAs, in ascending observing-time order. Matches populate the +associated Merges tables; unmatched sources become new AstroObjects +records. The resulting records are loaded into Merges_ and +AstroObjects_ tables. + +Matching runs in parallel by field, using multiple database-server cores, +and extends across field boundaries for sources near field edges. The +architecture can scale by moving the database server to a machine with +more cores and memory, as affordable. + +Lightcurve Statistics and Maintenance +===================================== + +A separate process computes lightcurve statistics after cross-matching +and stores them in AstroObjectsMeta_. Its script drops and +recreates all AstroObjectsMeta_ tables before computing and +inserting statistics, then explicitly vacuums and analyzes them at the end. + +Reprocessing generates new product versions, usually with the latest +designated best. Separate processes remove not-best lightcurve data +points from Sources and Merges_, then explicitly cluster, vacuum, +and analyze those tables. + +SOC-Sim Test Results +==================== + +The 7/22/2026 test with SOC sims covered 360 different fields, using a +match radius of 0.055 arcsec, or half a Roman WFI pixel: + +* ~250 million sources, regardless of ``flags`` value, were loaded into + Sources__ child tables in 1.3 hours with 8 parallel processes. +* ~198 million sources with ``flags = 0`` were cross-matched in + 35 minutes with 8 parallel processes. +* ~90 million AstroObjects records and 211,394,526 Merges records were + loaded into PostgreSQL. Of those merges (a.k.a. lightcurve data points), + 33,223 came from cross-matching across field boundaries, where the match + radius can extend across a boundary, increasing the merge count by 0.0157%. +* Computing statistics for ~90 million AstroObjects took 1.6 hours with + 8 parallel processes. diff --git a/docs/source/dev/database_connections.rst b/docs/source/dev/database_connections.rst index 226b9b58..980d8723 100644 --- a/docs/source/dev/database_connections.rst +++ b/docs/source/dev/database_connections.rst @@ -1,138 +1,116 @@ Database connections from pipeline code #################################################### -This page documents ``rapidpipe.db.connection``, the one connection path -pipeline code uses to reach the run-model tables. It replaces an earlier -narrative incident report carried on a predecessor branch (see "Two -incidents behind the design" below); that document does not appear in -this repository because it named private infrastructure, package -versions and work-item status that do not exist here. The design lessons -it recorded are kept, generalized, and traced to the code that -implements them. +``rapidpipe.db.connection`` is the one path pipeline code uses to connect +to the run-model tables. This page replaces a narrative incident report +from a predecessor branch, retaining its generalized lessons and tracing +them to the implementing code (see "Two incidents behind the design" +below). The original report is absent from this repository because it +named private infrastructure, package versions and work-item status that +do not exist here. What the connection module guarantees ************************************************ -**Where the endpoint and credentials come from.** ``connect()`` accepts -an explicit ``endpoint=`` (host, port, dbname) and ``credentials=`` -(user, password). A caller that already holds either -- for example -because it just read a parameter tree or resolved a secret under its own -role -- passes it directly. Anything not passed falls back to the -standard ``PG*`` environment variables (``PGHOST``, ``PGPORT``, -``PGDATABASE``, ``PGUSER``, ``PGPASSWORD``), and if ``RAPID_DB_SECRET_ID`` -is set in the environment and no explicit ``credentials=`` was given, the -credential is instead resolved from that AWS Secrets Manager secret via -``credentials_from_secret()``. A credential is never expected as a -command-line argument or read from a file committed to the repository. - -**Bounded connect retry.** Connecting retries on ``psycopg2.OperationalError`` -only, up to ``attempts`` times (default ``DEFAULT_CONNECT_ATTEMPTS = 5``), -with exponential backoff starting at ``backoff_initial`` -(``DEFAULT_BACKOFF_INITIAL_S = 0.5`` seconds), doubling by -``backoff_multiplier`` (``DEFAULT_BACKOFF_MULTIPLIER = 2.0``) each attempt -up to a cap of ``backoff_cap`` (``DEFAULT_BACKOFF_CAP_S = 8.0`` seconds). -Passing ``jitter=True`` replaces each wait with a random duration in -``[0, delay]`` so that many callers retrying off the same event do not -stay synchronized. Every one of these is an ordinary keyword argument a -caller can override. - -**TCP keepalives.** Every connection this module opens sets keepalive -parameters (``KEEPALIVES``, ``KEEPALIVES_IDLE_S``, -``KEEPALIVES_INTERVAL_S``, ``KEEPALIVES_COUNT``) and a -``TCP_USER_TIMEOUT_MS`` backstop, sized so a vanished peer is detected at -the socket in about a minute (30 seconds before the first probe, then up -to three probes ten seconds apart), rather than the kernel's own default -of two hours. +**Endpoint and credentials.** ``connect()`` accepts an explicit +``endpoint=`` (host, port, dbname) and ``credentials=`` (user, password). +Callers can pass either directly, for example after reading a parameter +tree or resolving a secret under their own role. Missing values fall back +to the standard ``PG*`` environment variables: ``PGHOST``, ``PGPORT``, +``PGDATABASE``, ``PGUSER`` and ``PGPASSWORD``. If ``RAPID_DB_SECRET_ID`` is +set in the environment and no explicit ``credentials=`` was given, +``credentials_from_secret()`` instead resolves the credential from that +AWS Secrets Manager secret. Credentials are never expected as command-line +arguments or read from files committed to the repository. + +**Bounded connect retry.** Only ``psycopg2.OperationalError`` triggers +connection retries, up to ``attempts`` times (default +``DEFAULT_CONNECT_ATTEMPTS = 5``). Exponential backoff starts at +``backoff_initial`` (``DEFAULT_BACKOFF_INITIAL_S = 0.5`` seconds), doubles +by ``backoff_multiplier`` (``DEFAULT_BACKOFF_MULTIPLIER = 2.0``) each +attempt, and stops growing at ``backoff_cap`` +(``DEFAULT_BACKOFF_CAP_S = 8.0`` seconds). ``jitter=True`` replaces each +wait with a random duration in ``[0, delay]`` to keep callers retrying +after the same event from staying synchronized. All are ordinary keyword +arguments callers can override. + +**TCP keepalives.** Every connection sets ``KEEPALIVES``, +``KEEPALIVES_IDLE_S``, ``KEEPALIVES_INTERVAL_S``, ``KEEPALIVES_COUNT`` and +a ``TCP_USER_TIMEOUT_MS`` backstop. These detect a vanished peer at the +socket in about a minute: 30 seconds before the first probe, then up to +three probes ten seconds apart, rather than the kernel's two-hour default. **Identification and timeout.** Every connection carries an -``application_name`` (truncated to PostgreSQL's 63-byte -``NAMEDATALEN - 1`` limit, visibly, before it reaches the server) and a -``connect_timeout`` (default ``DEFAULT_CONNECT_TIMEOUT_S = 10`` seconds) -bounding how long a single connection attempt can take. +``application_name``, visibly truncated to PostgreSQL's 63-byte +``NAMEDATALEN - 1`` limit before reaching the server. A ``connect_timeout`` +(default ``DEFAULT_CONNECT_TIMEOUT_S = 10`` seconds) bounds each +connection attempt. What it deliberately does not do ************************************************ -It does not reconnect and retry around a statement that was already in -flight when the connection died. The connect-retry described above -covers only the act of connecting; a connection that was established and -healthy, then closed mid-statement, is a server-side or pooler fault to -be diagnosed and fixed there, not papered over on the client. Hiding it -behind a silent retry would be actively harmful at the pipeline's -operating scale, where on the order of a thousand concurrent jobs may -each hold a connection: a pooler that is dropping payload connections -needs to be visible and fixed, not masked one retry at a time. - -It carries no pooler configuration, connection "lanes", or ``SET ROLE`` -role-widening bookkeeping. Pgbouncer (or any pooler) is server-side -infrastructure, provisioned and configured by ``rapid_systems``, per the -repository boundaries in the +Retries cover only connecting, never a statement in flight when a +connection dies. A healthy, established connection closed mid-statement +is a server-side or pooler fault to diagnose and fix there. Silent client +retries would harm the pipeline at its operating scale, where on the +order of a thousand concurrent jobs may each hold a connection, by +masking a pooler that drops payload connections. + +The module carries no pooler configuration, connection "lanes", or +``SET ROLE`` role-widening bookkeeping. Pgbouncer (or any pooler) is +server-side infrastructure provisioned and configured by +``rapid_systems``, per the repository boundaries in the `RAPID specification `_ -("Repositories"). This module only ever asks the pooler's, or the -database's, own listening port for one connection at a time. +("Repositories"). The module requests one connection at a time from the +pooler's or database's own listening port. Pooler configuration, including +per-user settings and timeouts, is owned and changed in ``rapid_systems``, +never in this repository. -The repository layer, ``rapidpipe.runs.repository``, opens no -transactions of its own beyond what this module provides: each of its -functions is documented as exactly one transaction per call, using -``transaction()`` below. +``rapidpipe.runs.repository`` opens no transactions beyond those this +module provides. Each repository function is documented as exactly one +transaction per call, using ``transaction()`` below. Two incidents behind the design ************************************************ -**August 2026.** A transaction-pooled pgbouncer instance began closing -the pipeline's freshly opened connections at ``age=0`` seconds with a -``client_idle_timeout`` closure, even though the pipeline's own database -user had no per-user timeout setting and the pooler's global -``client_idle_timeout`` was disabled. The cause turned out to be a -``client_idle_timeout`` value set on per-user configuration lines -belonging to a handful of other, human-operator database users; removing -those lines from the affected users stopped the closures immediately and -completely. That a per-user setting on unrelated users reached a user -with no per-user line of its own at all is recorded as unverified against -the pgbouncer issue tracker. The transferable lesson: a healthy -connection closed while a statement was in flight is a pooler defect, to -be diagnosed from the pooler's own log, not a reason to add -reconnect-and-retry on the client. The pipeline's fail-loud, nonzero exit -in the face of that closure was the correct behavior, and stayed -correct. +**August 2026.** A transaction-pooled pgbouncer instance closed freshly +opened pipeline connections at ``age=0`` seconds with a +``client_idle_timeout`` closure. The pipeline's database user had no +per-user timeout setting, and the pooler's global ``client_idle_timeout`` +was disabled. The cause was a ``client_idle_timeout`` value on per-user +configuration lines for a handful of other, human-operator database +users. Removing those lines stopped the closures immediately and +completely. That unrelated users' settings reached a user with no +per-user line remains unverified against the pgbouncer issue tracker. +The lesson: diagnose a healthy connection closed mid-statement as a +pooler defect from the pooler's own log, not by adding client +reconnect-and-retry. The pipeline's fail-loud, nonzero exit was and +remained correct. **September 2026.** A database host was replaced while a long-lived -client connection was still pointed at the address it had been using. -With no keepalives set, the client had no way to distinguish "quiet -connection" from "connection to a peer that no longer exists," and the -condition went undetected for hours, bounded only by the kernel's -default keepalive timeout. The lesson embodied in this module's -keepalive defaults: a client should set keepalive parameters on every -connection it opens so that a vanished peer is noticed at the socket in -about a minute. This is dead-peer detection at the socket level, wholly -distinct from statement-level retry; it says nothing about, and does -nothing for, a connection that a pooler actively closes while healthy, -which is the first incident above. +client connection still pointed at its old address. Without keepalives, +the client could not distinguish a quiet connection from a vanished peer. +The condition went undetected for hours, bounded only by the kernel's +default keepalive timeout. The lesson, embodied in this module's defaults: +set keepalives on every connection to detect a vanished peer at the socket +in about a minute. Socket-level dead-peer detection is distinct from +statement-level retry. It neither addresses nor helps with a pooler +actively closing a healthy connection, as in the first incident above. Diagnosing a dropped connection ************************************************ -#. The caller sees ``psycopg2.OperationalError`` (during connect, wrapped - by this module as ``ConnectionUnavailable`` once the retry budget is - exhausted) or a driver error surfaced mid-statement on an already - established connection. +* **Look first at the pooler's log.** Its close reason, such as + ``client_idle_timeout`` versus ``client close request``, is diagnostic + and not visible from the client. -#. Look first at the pooler's own log for the connection's close reason; - the reason string (for example ``client_idle_timeout`` versus - ``client close request``) is diagnostic and is not visible from the - client side. +* **Distinguish connection failure from mid-statement failure.** During + connect, ``psycopg2.OperationalError`` triggers automatic, bounded + retries with backoff. Once the budget is exhausted, ``connect()`` wraps + the error as ``ConnectionUnavailable``, naming the endpoint, user and + attempt count. A mid-statement closure on an established connection + instead propagates as a driver exception and is never retried. -#. A connect-time failure is retried automatically, bounded, with - backoff; if every attempt fails, ``connect()`` raises - ``ConnectionUnavailable`` naming the endpoint, user, and attempt - count. - -#. A close of an already-open connection, in the middle of a statement, - is never retried by this module; it propagates as a driver exception. - -#. Ruling out a vanished peer is a keepalive question, not a pooler-log - question: with the defaults above, a truly vanished peer is detected - within roughly a minute, so a hang longer than that points elsewhere. - -#. Pooler configuration, including per-user settings and timeouts, is - owned and changed in ``rapid_systems``, never in this repository. +* **Rule out a vanished peer through keepalives, not pooler logs.** The + defaults detect a vanished peer within roughly a minute; a longer hang + points elsewhere. diff --git a/docs/source/dev/notes.rst b/docs/source/dev/notes.rst index a2f4855d..861ad1b4 100644 --- a/docs/source/dev/notes.rst +++ b/docs/source/dev/notes.rst @@ -4,64 +4,57 @@ RAPID Pipeline Development Increasing AWS Cloud Limits ************************************ -Submit a ticket to the IPAC Support Group (ISG) requesting an AWS increase -in the relevant limit for the RAPID project -(this involves Wendy submitting a ticket to AWS). +Submit a ticket to the IPAC Support Group (ISG) requesting an increase in +the relevant AWS limit for RAPID. Wendy then submits a ticket to AWS. `ISG Request URL `_ -Login with your IPAC credentials (not sure whether VPN must be running). +Log in with your IPAC credentials; whether VPN must be running is uncertain. Development Guidelines ************************************ -#. Set up your text editor to clip trailing spaces when saving source-code file - (e.g., BBEdit has a setting that does this). +#. Configure your editor to remove trailing spaces on save and use spaces, + never tabs, for Python indentation. BBEdit has settings for both. -#. Ensure no tab characters are used for indentation in your Python code; use spaces always - (e.g., BBEdit has a setting that does this). +#. Keep revision diffs clear and unambiguous. Put extensive stylistic + changes in a separate revision so they do not hide behavior changes. -#. Think strategically when pushing a source-code file to the git repo whether a simple git diff between revisions - will allow a clear and unambiguous indication of the code changes. For example, numerous stylistic changes can - hide substantive changes that affect code behavior and should be deferred to a separate revision. +#. Before committing changes to someone else's code, establish the expected + level of trust and tell the author what to expect. -#. Before checking into the git repo modifications to someone else's source code, - let that person know what to expect (and assure there is the expected level of trust beforehand). +#. Write descriptive, self-explanatory commit messages so reports do not + require rereading the source code. -#. Your git commits should have self-explanatory descriptive messages (saves time not having to review source code later for reports). +#. Test changes before putting code into operations. Development is not + complete until the changes have been tested. -#. Always test code changes before the code is put into operations; the development is not done until - the code changes have been tested. +#. Include enough source-code comments. -#. Include a sufficiency of comments in your source code! +#. Run ``git pull`` often and before every ``git push`` to keep your RAPID + git repo up to date. -#. Remember to ``git pull`` before any ``git push`` and often, in order to make sure your RAPID git repo is up to date. +Exitcodes follow the Spitzer convention: -#. For exitcodes, we follow the Spitzer convention: - -============== ================ +============== ================================= Exitcode range Definition -============== ================ +============== ================================= [0,31] Normal termination, with messages [32,61] Warnings [64+] Error -============== ================ +============== ================================= GitHub Merging, Branching, and Pull Requests ******************************************** -This section describes the recommended git workflow for contributing to the -RAPID code base. - .. note:: - Pending approval of the team, we will be migrating to a ``dev`` branch - workflow and **disabling direct pushes to** ``main``. Once this is in - effect, all routine development will target ``dev``, and changes will - reach ``main`` only through pull requests. The instructions - below assume this workflow. + Migration to a ``dev`` branch workflow and **disabling direct pushes to** + ``main`` are pending team approval. Once in effect, routine development + will target ``dev``, and changes will reach ``main`` only through pull + requests. The recommended contribution workflow below assumes this model. GitHub Branches @@ -69,15 +62,14 @@ GitHub Branches The RAPID repository follows a two-branch model: -* ``main`` — the stable, production branch. Direct pushes will be disabled; +* ``main``: the stable, production branch. Direct pushes will be disabled; it is updated only via approved pull requests. -* ``dev`` — the active development branch. Day-to-day work lands here. +* ``dev``: the active development branch. Day-to-day work lands here. -The general rule of thumb: **small changes can go straight to** ``dev``, while -**large changes or new features get their own branch** off ``dev`` and can be -merged back via a pull request. The diagram below illustrates the full flow: -a feature branch off ``dev``, two commits of work, a pull request merging the -feature back into ``dev``, and ``dev`` later merging into ``main``. +**Small changes can go straight to** ``dev``; **large changes or new features +get their own branch** off ``dev`` and can merge back through a pull request. +The diagram shows a feature branch off ``dev``, two commits, a pull request +back into ``dev``, and a later merge of ``dev`` into ``main``. .. figure:: code_astro_feature_graph.png :width: 600 @@ -98,8 +90,8 @@ pushed directly to ``dev``. The basic cycle is **pull, commit, push**: git commit -m "Describe your change" git push -If you have unsaved changes and ``git pull`` reports a -conflict, stash your changes, pull, then re-apply your stash: +If ``git pull`` reports a conflict with unsaved changes, stash them, pull, +then re-apply the stash: .. code-block:: bash @@ -107,18 +99,18 @@ conflict, stash your changes, pull, then re-apply your stash: git pull git stash pop -After ``git stash pop``, resolve any conflicts that git reports -(see merge conflicts, below), then commit and push as above. +After ``git stash pop``, resolve any conflicts (see Resolving Merge +Conflicts below), then commit and push as above. -If you have a local commit that conflicts with a pulled commit, causing -``git pull`` to fail: +If ``git pull`` fails because a local commit conflicts with a pulled commit: .. code-block:: bash git pull --rebase -This will move HEAD to the latest commit from the remote branch and replay -your changes on top. Resolve the merge conflict (see below), and run: +This moves HEAD to the remote branch's latest commit and replays your +changes on top. Resolve the conflict (see Resolving Merge Conflicts below), +then run: .. code-block:: bash @@ -128,8 +120,8 @@ your changes on top. Resolve the merge conflict (see below), and run: Large Changes / Feature Additions ============================================ -For larger changes or new features, create a dedicated branch off ``dev`` -so that work-in-progress does not destabilize the shared branch. +Use a dedicated branch off ``dev`` for larger changes or new features to +keep work-in-progress from destabilizing the shared branch. Create a branch from ``dev`` @@ -143,26 +135,26 @@ If you are already on ``dev``, create and switch to a new branch: # or git checkout -b my_branch dev # if you are on another branch -Then push the branch to GitHub and set it to track the remote, so that -future ``git push`` / ``git pull`` commands work without extra arguments: +Push to GitHub and enable remote tracking so future ``git push`` / +``git pull`` commands need no extra arguments: .. code-block:: bash git push -u origin my_branch -After this, make commits as normal to your new branch. +Commit to the new branch as usual. Open a Pull Request back to ``dev`` -------------------------------------------- -When you are done with your feature branch or have completed major changes, -open a pull request on GitHub to merge it into ``dev``: +When the feature or major changes are complete, open a GitHub pull request +to merge the branch into ``dev``: 1. Push your latest commits (``git push``). -2. On GitHub, navigate to the repository. A banner usually appears offering - to **Compare & pull request** for your recently pushed branch — click it. - Otherwise, go to the **Pull requests** tab and click **New pull request**. +2. In the GitHub repository, click the **Compare & pull request** banner + that usually appears for a recently pushed branch. Otherwise, open + **Pull requests** and click **New pull request**. .. image:: pull_request_open.png :width: 600 @@ -190,10 +182,9 @@ open a pull request on GitHub to merge it into ``dev``: Close the branch after merging (optional) -------------------------------------------- -Once the pull request is merged, if you are finished editing a particular -feature, delete the branch to keep the repository tidy. On GitHub, click the -**Delete branch** button shown on the merged pull request. To delete the -branch locally and on the remote from the command line: +After merging, delete the branch if work on the feature is finished. +On GitHub, click **Delete branch** on the merged pull request. To delete +it locally and remotely from the command line: .. code-block:: bash @@ -230,15 +221,15 @@ Resolve any conflicts git reports, then commit the merge and push: Resolving Merge Conflicts ============================================ -A conflict happens when two changes touch the same lines of a file and git -cannot decide which to keep. This can come up after any of the operations -above. Git will report which files conflicted, for example:: +A conflict occurs when changes touch the same lines and git cannot choose +which to keep. Any operation above can cause one. Git reports the affected +files, for example:: Auto-merging pipeline.py CONFLICT (content): Merge conflict in pipeline.py Automatic merge failed; fix conflicts and then commit the result. -You can always list the files that still need attention: +List files that still need attention: .. code-block:: bash @@ -250,8 +241,7 @@ Conflicted files are shown under **"Unmerged paths"**. Editing the conflict markers -------------------------------------------- -Open each conflicted file. Git inserts markers around the disagreeing -sections: +Open each conflicted file and find git's conflict markers: .. code-block:: text @@ -261,14 +251,14 @@ sections: the incoming version of the lines >>>>>>> origin/dev -The block above ``=======`` is your current branch's version (``HEAD``); -the block below is the incoming version (here, ``origin/dev``). Edit the -file so it contains exactly what you want the final result to be, and -**delete all three marker lines** (``<<<<<<<``, ``=======``, ``>>>>>>>``). +Above ``=======`` is your current branch's version (``HEAD``); below it is +the incoming version (here, ``origin/dev``). Edit the file to the desired +result and **delete all three marker lines** (``<<<<<<<``, ``=======``, +``>>>>>>>``). .. note:: - VS Code makes this easier: it highlights each conflict and offers + VS Code highlights conflicts and offers **Accept Current Change**, **Accept Incoming Change**, **Accept Both Changes**, or **Compare Changes** buttons directly above the conflict. Click the one you want, or edit manually, then save the file. @@ -277,8 +267,7 @@ file so it contains exactly what you want the final result to be, and Completing the merge -------------------------------------------- -Once a file looks correct, stage it to mark the conflict resolved, then -repeat for every conflicted file: +Stage each corrected file to mark its conflict resolved: .. code-block:: bash @@ -305,24 +294,25 @@ Then push as usual. Bailing out -------------------------------------------- -If things get tangled and you want to start over, you can abort and return -to the state before the operation began: +To start over, abort and return to the state before the operation: .. code-block:: bash git merge --abort # during a conflicted merge git rebase --abort # during a conflicted rebase -If you applied a stash with ``git stash pop`` and want to undo it, note that -``pop`` removes the stash once applied; use ``git stash apply`` instead when -you want to keep the stash entry around as a safety net while resolving. +If you want to undo a stash applied with ``git stash pop``, remember that +``pop`` removes it once applied. Use ``git stash apply`` instead to keep the +stash as a safety net while resolving conflicts. Log into EC2 Instance Machine ******************************************** -This assumes you have already set up an EC2 instance under the AWS console, and that the EC2 instance is stopped. -Also, a key pair has been assigned to the EC2 instance, and the private key is installed in a ``.pem`` file on your laptop. +Start with a stopped EC2 instance already set up in the AWS console, an +assigned key pair, and its private key in a ``.pem`` file on your laptop. +The instance needs enough boot-disk space for ``docker build``; at least +32 GB is recommended. 1. Ensure the following environment variables are set on your laptop: @@ -335,9 +325,7 @@ Also, a key pair has been assigned to the EC2 instance, and the private key is i AWS_EC2_VOLUME_ID AWS_EC2_VOLUME_DEVICE -The two latter ones are only needed if your EC2 instance is to have an EBS volume attached. - -Your EC2 instance should have a large enough book-disk volume as ``docker build`` requires a lot of space; at least 32 GB is recommended. +The last two variables are needed only when attaching an EBS volume. 2. Ensure python3 is installed on your laptop and restart your EC2 instance: @@ -345,7 +333,7 @@ Your EC2 instance should have a large enough book-disk volume as ``docker build` python /source-code/location/rapid/aws/start_ec2_instance.py -Here is how to stop your EC2 instance later: +To stop the instance later: .. code-block:: @@ -361,10 +349,8 @@ Here is how to stop your EC2 instance later: Build Docker Image for RAPID Science Pipeline ********************************************* -Check your latest source-code changes into the RAPID git repo. - -Under root on your EC2 instance, check out the latest source code from the RAPID git repo, -and then build the Docker image for the RAPID pipeline: +Check your latest source-code changes into the RAPID git repo, then fetch +the latest code as root on your EC2 instance: .. code-block:: @@ -372,32 +358,32 @@ and then build the Docker image for the RAPID pipeline: cd /home/ubuntu/rapid git pull -The following command removes ALL Docker images from your EC2 instance, -but has the advantage of removing all Docker debris from the boot-disk volume, -thus reclaiming disk space: - -.. code-block:: - - docker system prune -a -f - .. warning:: - The above ``docker system prune`` command and the ``docker build`` command below will not work properly or as intended, - meaning the expected disk space will not be reclaimed, - unless all containers running the Docker image ``rapid_science_pipeline:1.0`` are stopped! + Stop all containers running ``rapid_science_pipeline:1.0`` before the + ``docker system prune`` and ``docker build`` commands below. Otherwise, + the commands will not work as intended and will not reclaim the expected + disk space. -Here is how to get a listing of your Docker containers that are running: +List running Docker containers: .. code-block:: docker ps -Here is how to get a listing of your Docker images: +List Docker images: .. code-block:: docker image ls +Remove ALL Docker images and debris from the instance's boot-disk volume +to reclaim space: + +.. code-block:: + + docker system prune -a -f + Rebuild the Docker image from scratch: .. code-block:: @@ -406,36 +392,37 @@ Rebuild the Docker image from scratch: docker build --build-arg RAPID_BRANCH= --file /home/ubuntu/rapid/docker/Dockerfile_ubuntu_runSingleSciencePipeline --tag rapid_science_pipeline:1.0 . -Push Docker image to the Amazon public elastic container registry (ECR): +Push to Amazon public elastic container registry (ECR) +====================================================== -Note that the RAPID-pipeline image has already been registered at +The RAPID-pipeline image is already registered at: .. code-block:: public.ecr.aws//rapid_science_pipeline -and so this step involves simply updating the Docker image in the registry. +This step updates that registry image. -Authenticate your Docker client to the registry as follows: +Authenticate your Docker client: .. code-block:: aws ecr-public get-login-password --region us-east-1 | docker login --username AWS --password-stdin public.ecr.aws/ -Now get the Docker image ID as follows: +Get the Docker image ID: .. code-block:: docker image ls -The response will be something like: +Example response: .. code-block:: REPOSITORY TAG IMAGE ID CREATED SIZE rapid_science_pipeline 1.0 a76b1373bfe2 6 minutes ago 2.36GB -Tag the Docker image with "latest" and push to ECR with these two commands: +Tag the image with "latest" and push to ECR with these two commands: .. code-block:: @@ -446,12 +433,13 @@ Tag the Docker image with "latest" and push to ECR with these two commands: Running an Instance of the RAPID Science Pipeline under AWS Batch ***************************************************************** -The following shows commands to launch an instance of the RAPID science pipeline as AWS Batch job. -The to-be-run-under-AWS-Batch Docker container rapid_science_pipeline:1.0 has /code built in, -so there is no need to mount an external volume for /code. -The container name is arbitrary, and is set to "russ-test-jobsubmit" in the example below. -Since this Docker image contains the ENTRYPOINT instruction, you must override it with the ``--entrypoint bash`` option -(and do not put ``bash`` at the end of the command). +Launch the RAPID science pipeline as an AWS Batch job with the commands +below. The Docker container rapid_science_pipeline:1.0 includes /code, so +no external volume is needed for /code. Its name is arbitrary; this example +uses "russ-test-jobsubmit". Override the image's ENTRYPOINT instruction +with ``--entrypoint bash``; do not put ``bash`` at the end of the command. + +Python 3.11 is required and installed in the image at /usr/bin/python3.11. .. code-block:: @@ -490,9 +478,10 @@ Since this Docker image contains the ENTRYPOINT instruction, you must override i exit -Python 3.11 is required and it is installed inside the Docker image (/usr/bin/python3.11). +Examine outputs +============================================ -After the AWS Batch job finishes, there are files written to S3 buckets that can be examined: +After the AWS Batch job finishes, examine the files written to S3 buckets: .. code-block:: @@ -560,28 +549,29 @@ After the AWS Batch job finishes, there are files written to S3 buckets that can 2025-03-14 09:19:57 730 20250314/jid1/refiminputs/refimage_unc_inputs.txt 2025-03-14 11:28:32 66890880 20250314/jid1/scorrimage_masked.fits -The general scheme for how the output files are organized in the S3 buckets is according to -processing date (Pacific Time) and the associated job ID. The same job ID can exist under -different processing dates if reprocessing occurred on different dates (reprocessing on the same date will overwrite products). +S3 output files are organized by processing date (Pacific Time) and job ID. +Reprocessing on different dates can place the same job ID under multiple +dates; reprocessing on the same date overwrites products. -The files under ``refiminputs`` are only written if the ``upload_inputs`` flag in the software is set to True. These are for -off-line analysis and rerunning awaicgen for experimental and tuning purposes. +Files under ``refiminputs`` are written only when the software's +``upload_inputs`` flag is True. They support off-line analysis and rerunning +awaicgen for experiments and tuning. -The reference-image products from ``awaicgen`` -are initially given generic filenames in these buckets, and, later, will be renamed to filenames like: +Reference-image products from ``awaicgen`` initially have generic filenames +in these buckets. After registration in the RAPID pipeline operations +database, they are renamed to filenames such as: .. code-block:: rapid_field1234567_fid7_ppid15_v2_rfid12394758_refimage.fits rapid_field1234567_fid7_ppid15_v2_rfid12394758_covmap.fits -The above filenames are created after these products are registered in the RAPID pipeline operations database. -The products are then copied to -a more permanent location (and ultimately archived in MAST). The ``ppid`` gives the pipeline number -that generated the reference image, which could be either the difference-image pipeline (``ppid=15``) -or a dedicated reference-image pipeline (``ppid=12``). +The products are then copied to a more permanent location and ultimately +archived in MAST. The ``ppid`` identifies the pipeline that generated the +reference image: either the difference-image pipeline (``ppid=15``) or a +dedicated reference-image pipeline (``ppid=12``). -Download and examine log file: +Download and examine the log file: .. code-block:: @@ -589,4 +579,3 @@ Download and examine log file: cat rapid_pipeline_job_20250314_jid1_log.txt Last modified: Tue 2026 Jun 16 8:48 a.m. -