Skip to content

Bump staging to 427548 - #856

Merged
anna-parker merged 1 commit into
mainfrom
create-pull-request/patch
Sep 28, 2026
Merged

anna-parker merged 1 commit into
mainfrom
create-pull-request/patch

Conversation

@github-actions

@github-actions github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

@anna-parker wants to update the staging deployment from pathoplexus/pathoplexus@b1a6c9b to pathoplexus/pathoplexus@4275489.

Commit date: 2026-09-28T15:31:06Z

Changelog

Comparison

pathoplexus/pathoplexus@b1a6c9b...4275489

Pre-merge checklists

Before merging this PR and bumping the staging version, clone the production database to staging and confirm this was successful:

  • DB clone: ssh bastion "cd pathoplexus/scripts/db-clone/ && ./clone-prod-to-staging.sh"
  • Restart staging backend: kubectl rollout restart deployment/loculus-backend -n staging
  • Run regressions tests to confirm staging is now identical to production (you will need to wait for SILO pods to import any new data)

3 ebola bdbv sequences were revised inbetween db clone steps, still seems safe to merge :-)

Post-merge checklists

After merging this PR, work through these checks (or provide justification for anything that was skipped) to confirm functionality and data integrity of Pathoplexus under the new version.

Standard checklist:

  • Run regression tests, highlight any observed differences and discuss why they are expected. If the new version includes new preprocessing pipeline versions: wait for reprocessing to finish and for all new SILO pods to come up first

There were sadly multiple new sequences ingested last evening, leading to a larger diff than I would have hoped for! In summary:

  • a new annotations column exists for all organisms
  • 3 revised ebola bdbv sequences (as before)
  • 31 newly ingested measles sequences (ingested on prod and staging)
  • 7 newly ingested rsv-a and 2 newly ingested rsv-b (ingested on prod and staging)
  • 2 newly ingested west-nile (ingested on prod and staging)
  • 2 newly ingested yf (ingested on prod and staging)

Fixes:

  • 39 newly ingested zika (only on staging -> these had the invalid characters in the name!!!! (Nebenf##hr)
  • 8 newly ingested yf on staging (also due to invalid characters in name!!! authors: Lee Cynthia K, [. U. S. ].; Monath Thomas P, [. U. S. ].; Guertin Patrick M, [. U. S. ].; Hayman Edward G, [. U. S. ].)

Changes:

  • For 84 direct mpox sequence submissions with ncbiVirusName and ncbiVirusTaxId, these are now deleted (it seems like they were actually passed through even for non INSDC users)
accessionVersion: PP_0031VMR.2
    ncbiVirusName: "Monkeypox virus" => ""
    ncbiVirusTaxId: "10244" => ""
  • Confirm submit/revise/revoke flow works (also pay close attention to the UI, document anything out of the ordinary):
    • Submit test data to the staging preview
      • Confirm submission goes through (paste AccessionVersion(s) here)
    • Revise an existing sequence (PP_007YJ7Y.2)
      • Confirm revision goes through
    • Revoke a test sequence (paste AccessionVersion here)
      • Confirm revocation goes through
image

Rollout-specific checklist:

Document any checks that were performed to test newly added features, database surgeries, etc. here.

@anna-parker
anna-parker merged commit 207c6bb into main Sep 28, 2026
1 check passed
@anna-parker
anna-parker deleted the create-pull-request/patch branch September 28, 2026 15:59
@anna-parker

anna-parker commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

As the changes introduced in loculus-project/loculus#7312 mean that now any direct submissions with noInput fields error we need to do a db surgery to rollout these changes.

Summary of Issue

Our PPX noInput fields:

yq '[.. | select(tag == "!!map" and .noInput == true) | .name] | unique | .[]' loculus_values/values.yaml
Details
sampleCollectionDateRangeLower
sampleCollectionDateRangeUpper
displayName
ncbiReleaseDate
earliestReleaseDate
ncbiUpdateDate
ncbiSubmitterCountry
insdcAccessionBase
insdcVersion
insdcAccessionFull
gcaAccession
length
hostNameScientific
hostNameCommon
hostTaxonId
ncbiSourceDb
ncbiVirusName
ncbiVirusTaxId
totalSnps
totalInsertedNucs
totalDeletedNucs
totalAmbiguousNucs
totalUnknownNucs
totalFrameShifts
frameShifts
completeness
totalStopCodons
stopCodons
outbreak
mutationsFromOutbreakFounder
lineage
serotype
lineage_S
clade
outbreakLineage
subtype
genotype

In total we have 10853 submissions with at least one of these fields (sql query below), (only 2630 without  'hostNameScientific', 'hostNameCommon').

Details
SELECT *
FROM (
    SELECT
        public.sequence_entries.*,
        (
            SELECT jsonb_object_agg(m.key, m.value)
            FROM jsonb_each(submitted_data->'metadata') AS m(key, value)
            WHERE m.key = ANY (ARRAY[
                'sampleCollectionDateRangeLower',
                'sampleCollectionDateRangeUpper',
                'displayName',
                'ncbiReleaseDate',
                'earliestReleaseDate',
                'ncbiUpdateDate',
                'ncbiSubmitterCountry',
                'insdcAccessionBase',
                'insdcVersion',
                'insdcAccessionFull',
                'gcaAccession',
                'length',
                'hostNameScientific',
                'hostNameCommon',
                'hostTaxonId',
                'ncbiSourceDb',
                'ncbiVirusName',
                'ncbiVirusTaxId',
                'totalSnps',
                'totalInsertedNucs',
                'totalDeletedNucs',
                'totalAmbiguousNucs',
                'totalUnknownNucs',
                'totalFrameShifts',
                'frameShifts',
                'completeness',
                'totalStopCodons',
                'stopCodons',
                'outbreak',
                'mutationsFromOutbreakFounder',
                'lineage',
                'serotype',
                'lineage_S',
                'clade',
                'outbreakLineage',
                'subtype',
                'genotype'
            ])
            AND m.value <> 'null'::jsonb
            AND m.value <> '""'::jsonb
        ) AS matched_metadata
    FROM public.sequence_entries
    WHERE group_id != 1
      AND is_revocation IS NOT TRUE
) AS q
WHERE matched_metadata IS NOT NULL;

Command to view the processing errors in the db:

Details
SELECT
    se.organism,
    se.accession,
    se.version,
    old_p.pipeline_version AS old_pipeline_version,
    new_p.pipeline_version AS new_pipeline_version,
    new_p.errors
FROM public.sequence_entries se
JOIN public.current_processing_pipeline cpp
  ON cpp.organism = se.organism
JOIN public.sequence_entries_preprocessed_data old_p
  ON old_p.accession = se.accession
 AND old_p.version = se.version
 AND old_p.pipeline_version = cpp.version
JOIN public.sequence_entries_preprocessed_data new_p
  ON new_p.accession = se.accession
 AND new_p.version = se.version
 AND new_p.pipeline_version = cpp.version + 1
WHERE old_p.errors = '[]'::jsonb
  AND new_p.errors <> '[]'::jsonb;

The odd one out was PP_001TG0W that has errors but is an INSDC ingest sequence, there are also 40 unreleased sequences (that have processing errors in both versions). The PP_001TG0W processing error was a temporary AWS S3 upload error (see slack thread for details: https://loculus.slack.com/archives/C05G172HL6L/p1790611195852959) - we should add a retry to the code for such cases.

@anna-parker

anna-parker commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

DB surgery steps

  1. Update the submitted_data, removing the noInput fields (they are set to empty string), note the archive_of_submitted_data stores the original input:
Details
UPDATE public.sequence_entries AS se
SET submitted_data = jsonb_set(
    se.submitted_data,
    '{metadata}',
    (
        SELECT jsonb_object_agg(
            m.key,
            CASE
                WHEN m.key = ANY (ARRAY[
                    'sampleCollectionDateRangeLower',
                    'sampleCollectionDateRangeUpper',
                    'displayName',
                    'ncbiReleaseDate',
                    'earliestReleaseDate',
                    'ncbiUpdateDate',
                    'ncbiSubmitterCountry',
                    'insdcAccessionBase',
                    'insdcVersion',
                    'insdcAccessionFull',
                    'gcaAccession',
                    'length',
                    'hostNameScientific',
                    'hostNameCommon',
                    'hostTaxonId',
                    'ncbiSourceDb',
                    'ncbiVirusName',
                    'ncbiVirusTaxId',
                    'totalSnps',
                    'totalInsertedNucs',
                    'totalDeletedNucs',
                    'totalAmbiguousNucs',
                    'totalUnknownNucs',
                    'totalFrameShifts',
                    'frameShifts',
                    'completeness',
                    'totalStopCodons',
                    'stopCodons',
                    'outbreak',
                    'mutationsFromOutbreakFounder',
                    'lineage',
                    'serotype',
                    'lineage_S',
                    'clade',
                    'outbreakLineage',
                    'subtype',
                    'genotype'
                ])
                AND m.value <> 'null'::jsonb
                AND m.value <> '""'::jsonb
                THEN '""'::jsonb
                ELSE m.value
            END
        )
        FROM jsonb_each(se.submitted_data->'metadata') AS m(key, value)
    )
)
WHERE se.group_id != 1
  AND se.is_revocation IS NOT TRUE
  AND EXISTS (
      SELECT 1
      FROM jsonb_each(se.submitted_data->'metadata') AS m(key, value)
      WHERE m.key = ANY (ARRAY[
          'sampleCollectionDateRangeLower',
          'sampleCollectionDateRangeUpper',
          'displayName',
          'ncbiReleaseDate',
          'earliestReleaseDate',
          'ncbiUpdateDate',
          'ncbiSubmitterCountry',
          'insdcAccessionBase',
          'insdcVersion',
          'insdcAccessionFull',
          'gcaAccession',
          'length',
          'hostNameScientific',
          'hostNameCommon',
          'hostTaxonId',
          'ncbiSourceDb',
          'ncbiVirusName',
          'ncbiVirusTaxId',
          'totalSnps',
          'totalInsertedNucs',
          'totalDeletedNucs',
          'totalAmbiguousNucs',
          'totalUnknownNucs',
          'totalFrameShifts',
          'frameShifts',
          'completeness',
          'totalStopCodons',
          'stopCodons',
          'outbreak',
          'mutationsFromOutbreakFounder',
          'lineage',
          'serotype',
          'lineage_S',
          'clade',
          'outbreakLineage',
          'subtype',
          'genotype'
      ])
      AND m.value <> 'null'::jsonb
      AND m.value <> '""'::jsonb
  );
  1. Delete the processed_data entries that have errors, so that they can be reprocessed after the db surgery:
Details
BEGIN;

CREATE TEMP TABLE rows_to_delete ON COMMIT DROP AS
SELECT
    new_p.accession,
    new_p.version,
    new_p.pipeline_version
FROM public.sequence_entries se
JOIN public.current_processing_pipeline cpp
  ON cpp.organism = se.organism
JOIN public.sequence_entries_preprocessed_data old_p
  ON old_p.accession = se.accession
 AND old_p.version = se.version
 AND old_p.pipeline_version = cpp.version
JOIN public.sequence_entries_preprocessed_data new_p
  ON new_p.accession = se.accession
 AND new_p.version = se.version
 AND new_p.pipeline_version = cpp.version + 1
WHERE old_p.errors = '[]'::jsonb
  AND new_p.errors <> '[]'::jsonb;

ANALYZE rows_to_delete;

DELETE FROM public.sequence_entries_preprocessed_data p
USING rows_to_delete d
WHERE p.accession = d.accession
  AND p.version = d.version
  AND p.pipeline_version = d.pipeline_version;

COMMIT;
  1. Restart the prepro pods (a backend error causes the etag to only rely on the changes from the current pipeline version so it is not updated when there are other changes), leading the prepro pods to not get any sequences to reprocess (alternatively waiting 1h would work I think)

All have successfully reprocessed:
image

@anna-parker

Copy link
Copy Markdown
Contributor

I just realized that the errors for no input fields are actually really tedious, if I submit with these no input fields I am unable to remove them via the edit page as they do not show up... the only way is to fully delete the submission... I didnt test this flow through before so I didnt realize how poor the UX is... I think we cant release this

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants