DON'T MERGE: query engine review fixes preview (stacked on #7431) - #7435
Draft
corneliusroemer-agent wants to merge 167 commits into
Draft
corneliusroemer-agent wants to merge 167 commits into
corneliusroemer-agent wants to merge 167 commits into
Conversation
…nce_entries
The query projector reads each dirty batch with `sequence_entries_view ... where accession in (...)`.
Postgres cannot push an IN list into the view's `GROUP BY` subquery over external_metadata, so every
batch of two or more accessions aggregated the whole external_metadata table (all organisms).
New view sequence_entries_lateral_view (V1.38): same columns and rows, but external metadata is
aggregated per row with LEFT JOIN LATERAL (keeping the GROUP BY, so "no external metadata" stays NULL
rather than jsonb_merge_agg's '{}'). streamReleasedSubmissionsWithCompressedSequences uses it when
an accession filter is given; full streams (get-released-data, full rebuild) stay on
sequence_entries_view, whose plan is better for them. CREATE VIEW only takes AccessShare locks, so
the migration does not wait for running get-released-data streams (a CREATE OR REPLACE of the
existing view would need ACCESS EXCLUSIVE on it).
Measured on PG16 with 1M released entries and 1M external_metadata rows (in cache):
- 2000 random accessions: 250 ms -> 25 ms
- 10 newest accessions (the usual "just approved" batch): 230 ms -> 0.1 ms
- full organism stream: unchanged (still uses sequence_entries_view; through the lateral view it
would be ~15% slower, same md5 of the output)
The generic plan with bind parameters (as Exposed sends them) is the same nested loop.
SequenceEntriesLateralViewTest compares both views row by row (entries without, with one and with
two external metadata updaters, revised and revoked entries, unprocessed entries) and the
accession-filtered released data against the full stream, so the definitions cannot drift apart
silently.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s Hikari's default 10) Nothing configured the pool, so each backend replica ran Hikari's default of 10 connections. The query engine adds bulk readers to that pool: up to 6 connections per metadata export, 4 per sequence download, ~9 for the projector while it drains (4 writers, each also allocating ids on a second connection, plus the leader-lock connection) and 5 per organism while an index loads. On the preview, 8 concurrent metadata exports slowed every DB-backed request 4-5x. The pool size is now 30 (Helm: backendDatabasePoolSize), with minimum-idle 10 so an idle replica keeps as many connections open as before. Downloads fetch chunk by chunk and give the connection back after each chunk, so with 30 connections a waiting request waits for one chunk of another request, not for a whole download; the 30 s Hikari connectionTimeout is only reachable when most connections are held for long (e.g. index loads of several organisms at startup). Budget: replicas x 30 plus other clients must stay below Postgres max_connections (default 100), so 2 replicas (Pathoplexus) leave ~40 for everything else. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… at a slow rate A projection can go stale without any later change to its source rows: the one-by-one fallback drops an accession from the dirty queue after any error (transient ones included), and a trigger can be missed (e.g. writes with session_replication_role = replica). Nothing recomputed such entries except a full rebuild. Each projector run now marks the next released accessions of every organism dirty, in accession order with a per-organism cursor that wraps around, at loculus.query-engine.reconcile-accessions-per-second (default 50, 0 turns it off). One pass takes (released accessions / 50) s: ~21 min for Pathoplexus' largest organism (63k), ~5.5 h at 1M. Unchanged entries cost a read and a projection but no write (change-only upserts; sequence data skipped by source_hash), and the queue gains at most one batch per run, so live changes never wait behind a whole-organism sweep. The cursor query is an index range scan (0.2 ms for 50 accessions at 1M rows). The drain now logs only runs that changed something, as the reconcile marks accessions every run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing on submissions Inserting a dirty-queue key that an uncommitted submission transaction has also inserted waits for that transaction. A reconcile step that waited while holding its own freshly inserted keys could form the same deadlock cycle as a group rename during an approval. With lock_timeout = 100 ms (below deadlock_timeout, 1 s) the step is abandoned before a cycle can be detected, and retried from the same cursor in the next run. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… hours At 50 accessions/s per organism a small organism finished its reconcile pass within seconds and started over at once, so the reconcile re-projected it continuously and never idled. A new pass now starts at most every loculus.query-engine.reconcile-pass-interval-minutes (default 360) after the previous one started; large organisms (a pass longer than that, e.g. 1M at 50/s = 5.5 h) run back to back as before. After a restart the first pass starts immediately. The cursor query stays cheap with organisms interleaved in accession order (one accession sequence for all organisms), measured at 1M rows split over 5 organisms of 200k plus one of 10k interleaved at 1%: 0.1-0.4 ms per step for a 200k organism, 1.6-2 ms for the 1% one, 0.03 ms at the end of a pass; an organism clustered at the start of the key range gets the organism index (0.2-0.3 ms). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The engine wrote `null` for lineage nodes without parents or aliases. LAPIS
writes `{}`, and the website's lineageDefinition schema (types/lapis.ts)
requires an object per node, so the null roots failed validation and broke
lineage autocomplete.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Out-of-range values were accepted: -1 returned every mutation, 2 returned nothing. LAPIS (via SILO) answers 400 with "Error from SILO: Invalid proportion: minProportion must be in interval [0.0, 1.0]"; the engine now returns the same status and message. NaN is rejected as well. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Mutation and insertion requests over <=100 ids loaded the rows from Postgres and counted them directly. That costs per id (~0.04 ms locally, 0.2-0.4 ms measured on the preview's real rows) while the bitmap path is a flat 1-3 ms at 60k-1M entries, so from roughly 30-50 ids on the "fast path" was the slower one (preview: +20-40 ms at 100 ids). In-process probe, synthetic SARS-CoV-2-shaped rows, median ms (bitmap | rows from local Postgres): N=60k k=1 1.1 | 0.24 k=10 0.9 | 0.55 k=100 1.4 | 3.9 N=1M k=1 2.1 | 0.11 k=10 3.1 | 0.46 k=100 2.4 | 3.9 Rows stay much faster for a single sequence (the website's details page), so the limit drops to 20 instead of removing the path. The complement path (all but <=100 ids) keeps its own limit of 100. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
tick() caught only Exception. scheduleWithFixedDelay cancels all later runs of a task that throws and keeps the throwable in a discarded future, so any Error (StackOverflowError, OutOfMemoryError, ...) froze that organism's index for good without a log line. tick() now logs every Throwable at ERROR and keeps ticking, except OutOfMemoryError, which is logged and rethrown: the backend JVM options get -XX:+ExitOnOutOfMemoryError so an OOM anywhere (index load, a large request) restarts the pod instead of leaving it half-working. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…loaded The pod reported Ready as soon as Spring started, while the indexes were still building, so every rollout served 503 "initializing" to LAPIS requests routed to the new pod (18-30 s at 1M entries per the spec). A queryIndex health indicator is OUT_OF_SERVICE until every organism's index has finished its first load, and is added to the readiness group only: liveness and the startup probe are unaffected, so a slow build never restarts the pod. Later reloads keep serving the old index and do not touch readiness. The indicator is registered even with the query engine disabled (then always UP), since a health group must name an existing contributor. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The website's free-text and substring search sends `field.regex` as "(?i)" plus the backslash-escaped search text, a new pattern per search value. Each new pattern is evaluated with RE2J against every distinct value of the column, under the organism's read lock. Patterns that are an ASCII literal (optionally case-insensitive, with only escaped metacharacters) now use a plain substring search with RE2's ASCII case folding (including U+212A and U+017F); everything else still goes through RE2J. A randomized test checks agreement with RE2J. New-pattern cost on an accession-like column, median: 27k distinct values: 6.7 ms -> 1.1 ms 1M distinct values: 240 ms -> 36 ms Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A full reload (changelog backlog above max(50k, 10% of the index), e.g. after a projection rebuild) built the new index while the old one kept serving, so the heap held two copies of the organism's index plus loader buffers at the peak. The old index is now dropped first; that organism answers 503 until the reload finishes (1.5 s for 300k synthetic SARS-CoV-2-shaped entries here, 18-30 s at 1M per the spec). Smallest -Xmx at which a reload of a 276 MB index (300k entries) succeeds: before 620-640 MB, after <= 380 MB. Readiness now tracks the first load per organism instead of the presence of an index, since every replica reloads at the same time and gating readiness on it would take all of them out of service together. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each organism's tracker reloaded on its own thread, so a projection rebuild across organisms reloaded them all at once: one index rebuild plus ~160k parsed loader rows and 4 pool connections per organism. A service-wide lock now serialises reloads. The old index is dropped only once an organism's turn comes, so organisms waiting in the queue keep serving their current index. First loads at startup stay concurrent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ndJvmExtraOpts For the PR #7431 query-engine scale test on a Pathoplexus preview: the in-cluster development Postgres had a hardcoded 2Gi limit, stock settings and no volume, and the backend JVM flags could not be extended from values. All new keys default to the previous rendering. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…e query-engine preview Preview-only values for reviewing PR #7431 with the H/G fixes on a Loculus preview. - mpox ported from Pathoplexus (schema, preprocessing v28 with nextclade nextstrain/mpox/all-clades@2026-07-07--14-07-11Z, reference genomes) plus the mpoxOutbreakLineage definition, added to defaultOrganisms so Theo's default set stays. Ingest from NCBI taxon 3431483 with subsample_fraction 0.2 (~2.7k of 13.4k records). - Backend: 1 replica, 8Gi, MaxRAMPercentage=80 and GC logging via backendJvmExtraOpts. - Dev Postgres: 4Gi limit, max_connections=200, shared_buffers=1GB, 1Gi /dev/shm. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
Replicas tail query_changelog by seq and skip a gap after 10 s. With a global bigserial, a seq that commits more than 10 s late, or that is still uncommitted when fullLoad reads max(seq), was never applied to that replica's index, and the reconcile job does not heal it (an unchanged re-projection writes no changelog row). ProjectionWriter now takes the organism's query_engine_state row lock first and then inserts seq = max(seq of the organism) + row_number in a separate statement, so seqs of an organism become visible in commit order and without holes. V1.39 makes the primary key (organism, seq); the bigserial default stays for backends on older code during a rolling update. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…fails Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…ngest fits the 1800 s deadline Mpox parses the whole ~2.5 GB NCBI package before subsampling; Pathoplexus production does that within 1800 s on 1 CPU. The request stays at 50m. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
Pure re-indentation and line unwrapping from `yamlfmt` 0.21.0; the parsed YAML is identical. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…sts 3Gi so a rolling update can schedule A push that changed values recreated the dev Postgres empty while the old backend kept running against it (relation ... does not exist), and the new 8Gi-request backend pod stayed Pending next to the old one. The 8Gi limit is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
… backoff A reload that failed partway left the tracker with no index, and the next tick then ran a plain full load outside reloadLock. That broke the one-copy-at-a-time invariant the heap sizing relies on, just when something was already wrong, and retried every tail interval with a stack trace each time. Every load of an organism after its first now takes reloadLock; first loads at startup stay concurrent. A failed load waits 1 s before the next attempt, doubling up to 5 min, and the backoff resets on success. Tail failures keep retrying every tick, so a database blip does not delay release-to-queryable. Also corrects the stale pool-size note on LOAD_READERS. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
… other error The rethrow did not do what its comment said. A heap OOM never reaches the catch: -XX:+ExitOnOutOfMemoryError terminates the JVM at the failed allocation. An OOM thrown by Java code, such as "Required array length is too large" from collection growth, was caught and rethrown, which silently cancelled the organism's scheduled task while the JVM kept running: after a dropped index, that organism answered 503 for good with readiness UP. Such an error allocated nothing, so the heap is intact, and it is usually deterministic, so a restart would not fix it. Halting the JVM would turn it into a crash loop that takes submission down with a single replica, and would also kill a test JVM. It now goes through the Throwable branch: logged, and retried with the load backoff. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
Exposed's transaction {} retries every SQLException up to its default of 3 attempts, so a 100 ms lock timeout in
the reconcile marking cost up to 300 ms of the leader tick, each attempt logging Exposed's WARN with a stack trace.
The next projector run already retries the step from the same cursor, so the transaction now makes one attempt.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…thout a preflight The browser's conditional LAPIS cache set If-None-Match from JavaScript on every repeat. A header set from JS makes a cross-origin GET non-simple, so each repeat was preceded by an OPTIONS preflight: a full network round trip, and the preflight cache is per URL, so every new filter combination paid it. The browser's axios instance now caches POST only. GETs go straight to the network adapter (XHR), which uses the browser's HTTP cache: the query engine sends Cache-Control: no-cache with an ETag, so the browser stores the response and revalidates it with its own If-None-Match, which needs no preflight. XHR surfaces the revalidation's 304 as a 200 with the stored body, and nothing in the website relies on seeing a 304 or the x-loculus-lapis-cache header for a GET. POSTs are preflighted anyway (JSON content type) and no HTTP cache stores them, so they keep the JS cache. The server-side render cache is unchanged and still caches GET and POST. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…nly changed records Every run wrote the whole NCBI corpus to disk several times: seqkit's reformatted sequences.fasta, then sequences.ndjson, then the metadata TSVs and ndjson, and prepare_files read the full sequences.ndjson again. For SARS-CoV-2 that is ~85 GB written twice and read about four times, whether or not anything changed. For non-segmented organisms, select_changed_records replaces format_ncbi_metadata, filter_out_depositions, calculate_sequence_hashes, prepare_metadata, metadata_filter and compare_hashes. It reads genomic.fna once to hash each sequence and note where its record is, streams the data report in order through the same per-record functions the old scripts used (now factored out of them), and writes the compare_hashes outputs plus the metadata and sequences of the records to submit or revise. prepare_files runs unchanged on those. On 100k SARS-CoV-2 records with a routine mix of changes, every backend-facing output is byte-identical to the old pipeline, and the rules write 94 MB instead of 6.2 GB. get_previous_submissions no longer waits for the prepared metadata, so it runs alongside the download. Segmented organisms keep the old rules: grouping recomputes the hash of a group from all its segments, so records can't be dropped before grouping. construct_submitted_dict now streams the previous submissions and keeps two versions per Loculus accession instead of every entry (~1 kB per previous submission, GBs for SARS-CoV-2), with the same result, including which version wins a tie and the curated flag of revoked accessions. Both paths use it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…cords With mirror_format: zst, the fetch rule decompressed the mirror's genomic.fna.zst to disk (after seqkit grep for released_after) only for select_changed_records to read it once and seek back into it for the changed records. For SARS-CoV-2 that is an ~85 GB write and read per run, for a file that is 2.8 GB compressed. Non-segmented organisms without a minimizer now keep the download compressed and select_changed_records decompresses it itself (stdlib compression.zstd, ~6 GB/s here): once to hash, once more to extract the changed records by byte range, since seeking forward in the decompressed stream is as fast as reading it. The release filter moves from seqkit grep into the parser (--fasta-ids = released_accessions.txt): records outside it are skipped before their sequence is joined or decoded, so an invalid header outside the filter still does not fail the run. The filter ids are the index's keys, so they cost no extra memory. Segmented organisms, a configured minimizer and the other fetch paths still write genomic.fna, which nextclade and calculate_sequence_hashes read. Byte-identical to 98f8e18 on 100k SARS-CoV-2 records with the fetch rule run from a local mirror: routine, routine with released_after (31% kept), 50k bootstrap, and subsample + filter. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
bytes.find(b"\n>") costs ~0.7 s/GB on the wrapped NCBI FASTA, because every line starts a partial match; finding ">" is a memchr at ~0.02 s/GB. The difference was most of the per-GB cost of the records the release filter skips (for SARS-CoV-2, ~190 of the mirror's ~270 GB) and a quarter of select_changed_records on the records it keeps. 100k SARS-CoV-2 records: select_changed_records 10.1 -> 8.0 s (routine), 5.5 -> 3.6 s (routine with 31% released after the cut-off). Outputs byte-identical. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…ntry get_previous_submissions peaked at 6.5 GB for 2.84M SARS-CoV-2 submissions: get_submitted read the whole /get-submitted-metadata body (requests without stream), held every entry as a dict (~2 kB each with the body), then reopened the output file once per entry to append it. The response is now streamed (stream=True, 1 MB chunks) into a temporary file and counted as it goes; as before, it is retried until the count equals x-total-records. The statuses are still fetched after the entries, so that a version submitted in between has its status rather than UNKNOWN; the entries are then rewritten from the temporary file with their status through one file handle. That costs one extra write and read of the file (~0.5 GB at 2.84M), not memory. The output is byte-identical (mock backend at 100k and 50k entries, with an interrupted first stream, and with none). Peak RSS at 100k: 254 -> 168 MB; per entry 2.2 -> 1.3 kB. What is left is get_sequence_status parsing the whole /get-sequences body at once (~1.3 kB per entry transiently, ~380 B kept), which streaming this endpoint does not touch. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
With get_submitted streamed, get_sequence_status was the peak of get_previous_submissions:
response.json() built the list of every sequence entry (~1.3 kB each while parsing, ~3.7 GB at
2.84M SARS-CoV-2 entries) before the {accession: {version: status}} map was filled from it.
An object_hook now records each entry's status as the parser builds the entry and drops the
entry, and the map is keyed by "accession.version" with interned statuses (93 instead of 377 B per
entry). get_submitted is its only caller. On a decoding error the result is empty, as before.
Mock backend, previous_submissions.ndjson byte-identical at 100k and 50k, after a retried stream
and with no entries. Peak RSS of get_submitted at 100k: 168 -> 90 MB; slope 1.3 -> 0.6 kB per
entry, ~1.7 GB extrapolated to 2.84M (6.3 GB before streaming). Also close a 423 response before
retrying, now that GETs may stream.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
… preview (~3.83M) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The ingest calls /get-sequences only to learn the status of every previous submission, which it then joins onto the get-submitted-metadata entries. At 2.84M SARS-CoV-2 entries that second call is the ingest's memory peak and over a minute of wall time. With the status in the same response, the ingest can drop it. The status is the value of sequence_entries_view.status, but computed with a correlated subquery that reads sequence_entries_preprocessed_data only for entries that are neither released nor revocations. Selecting the view's status column instead makes Postgres hash-join the whole preprocessed-data table: on 1M synthetic entries that is +1.3-1.7 s and 145 MB of temp files per call, against +0.06 s for the subquery (4.43 s baseline). A test checks that every status matches what /get-sequences returns. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…ed-metadata The backend now returns the status with each entry, so get_submitted no longer calls /get-sequences to build an accession.version -> status map. That call parsed one JSON document of every entry and was the ingest's memory peak (2.40 GB live at 2.84M SARS-CoV-2 entries). The entries are written to disk as they stream in and moved into place, without the second rewrite pass that added the statuses. The status now comes from the same snapshot as the metadata. Before, the statuses were fetched after the entries, so that a version submitted in between got its status rather than UNKNOWN; that race is gone. status is written as the last key of each entry, where the files always had it, so previous_submissions.ndjson is byte-identical to before (checked against a mock backend at 100k and 50k entries, after a truncated stream, and for 0 entries). An entry without status (a backend older than this ingest) fails the run instead of being treated as unapproved. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
A routine SARS-CoV-2 ingest on the preview failed in get_previous_submissions with 502 Bad Gateway while a backend pod shut down during a rollout, and that failed the whole run. get_submitted_metadata now fetches the whole stream again, into a fresh file, after a connection error, a timeout, a 5xx, a stream cut mid-line or mid-chunk, a count that differs from x-total-records, or an entry without status (an old backend pod still serving during the rollout that brings the status field): up to 8 attempts, with 10 s doubling to at most 120 s between them (~8.5 min of backoff), then it fails. Each attempt can additionally wait on get_jwt's Keycloak retry (300 s) or the request timeout. Client errors (4xx) still fail at once. make_request fetches a new token for every attempt. A cut stream used to fail the run on the invalid last line, and a short count was retried every 60 s without end; both now take the same bounded backoff. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
autoapprove calls approve-processed-data every ~70 s per organism. It selected PROCESSED entries from sequence_entries_view, whose status is a CASE over the join with sequence_entries_preprocessed_data, so Postgres joined the organism's whole table on every call: 6.9 s and ~67 MB of temp files per call for 3.8M SARS-CoV-2 entries with nothing to approve (61 % of the preview DB's temp writes). Every PROCESSED entry is unreleased, so filtering on released_at IS NULL changes no result. With a partial index on unreleased entries the plan starts from the backlog: 1,390 -> 12 ms on 750k entries locally. Without the index Postgres misestimates the predicate (438k rows expected, 1,206 actual) and still scans the preprocessed data. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The default 200 held autovacuum to 15-34 MB/s on the preview's NVMe; one TOAST vacuum took 352 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The released_at IS NULL predicate on the view (5d7a358) is planned from table statistics. After a large ingest they are stale: at 3.8M SARS-CoV-2 entries with 1,206 unreleased, Postgres expected 416k and still hash-joined all of sequence_entries_preprocessed_data (2.4 s per autoapprove call). Without an accession filter, approve now reads the organism's unreleased entries from sequence_entries_unreleased_idx first and restricts the view query to those pairs, bound as two unnest arrays. That is planned as a nested loop of primary-key lookups regardless of statistics: 4 ms to plan and 32 ms to run on the same data (an IN list of row literals took 0.4 s to plan for 1,200 pairs). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The preview gets all 16 PPX organisms (pathoplexus d58747e, incl. chikungunya) with PPX's own ingest, preprocessing, EMBL annotation files and ENA deposition config, so the query engine, preprocessing and ingest see a PPX-shaped deployment next to SARS-CoV-2. sars-cov-2, cchf-multi-ref and enteroviruses are copied resolved from the chart's defaultOrganisms and render identically (helm template diff: only the values-hash annotations and the ingest schedules, which are spread over the organism count, change). The pure test organisms (dummy-organism, dummy-organism-with-files, not-aligned-organism) are dropped; they hold no data. Only each PPX organism's newest pipeline version is kept: PPX is mid-reprocess into it, and a fresh organism would otherwise be processed twice. mpox, cchf, west-nile and ebola-sudan move from the upstream config to PPX's and are reprocessed (~26k entries; mpox gains EMBL files). Memory requests for the PPX preprocessing and ingest pods are set from their peak working set on PPX staging during its 2026-09-28 reprocess, so the scheduler sees the bootstrap's real footprint. Generated by investigations/2026-09-28-loculus-pr7431-query-engine/59-data/port_ppx_organisms.py. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…ine, with small pods Follow-up to f6fc913ad: - No new EMBL annotation files anywhere (create_embl_file: false for every organism). The annotations category stays only where entries already carry files (cchf, west-nile, ebola-sudan and the unchanged chart organisms), so their links still resolve. - mpox, cchf, west-nile and ebola-sudan keep the pipeline versions they run on the preview (28, 1, 1, 1+2) with PPX's config, so none of them is reprocessed. Their reference sequences and genes are identical to PPX's; cchf's PPX-only lineage_S stays empty on existing entries until a later bump. - One preprocessing pod per PPX organism; SARS-CoV-2 preprocessing goes from 20 pods to 2 (config unchanged, the old count is in a comment next to it). - PPX organisms' preprocessing and ingest requests are near typical use (128Mi; mpox 256Mi / 1Gi ingest, rsv 512Mi ingest) and limits are the measured peak x 1.3, at least 512Mi (peaks from PPX staging/production and this preview). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
A revision adds a new version of an existing accession, which sorts before the claim cursor, so only the next full scan (every 10 s) found it. The revise endpoint now remembers the new versions after its transaction commits, and the organism's next claim takes them first. The memory is per backend replica and bounded; anything it misses is still found by the full scan, which can then run less often. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
… more often With the default scale factors (vacuum 20 %, analyze 10 %) sequence_entries went 12 h without an analyze after a 1M-entry ingest, and the stale statistics chose a 2.4 s plan over a 30 ms one for the approve query. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
… s on the preview New entries and revisions submitted to the same replica are claimed without a full scan, which is left to catch stale resets, rolled-back claims and revisions on other replicas. At 3.8M SARS-CoV-2 entries each full scan takes 3.6-3.9 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Nine of the Pathoplexus organism images exist only in the Pathoplexus website, so the preview's landing page and menu showed broken images. Chikungunya is on staging only, so its image comes from the pathoplexus repo at a fixed commit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…PPX main peak x1.3 Dengue preprocessing was OOM-killed at 768Mi; PPX main's 7-day per-pod peak is 1,118 MiB. Mpox's 3200Mi left 13% over its 2,834 MiB peak. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
A CPU limit only stops a container from using idle cores: under contention the request already sets its share. On the preview, ingest was throttled in 27-85% of CFS periods (cchf, rsv, mpox), stretching runs and holding their memory longer; init containers (config-processor, keycloak theme, flyway) were capped at 500m. Memory limits are unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Both nextclade runs were hard-coded to --jobs=1, so a pod used one core for alignment however many it had; the only way to scale was more replicas, each with its own Python process and loaded dataset. nextclade parallelises over the sequences of a batch, so nextclade_jobs: N uses N cores per pod. The default stays 1 (no behaviour change). The preview sets 4 for sars-cov-2, dengue and mpox, which have no CPU limit since 47c10e8. Portable-to-main: yes (config.py + nextclade.py; values change is preview-only) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…aiting for another write Deleting preprocessed rows makes their entries claimable again, but they sort before the claim cursor. The delete bumps the update tracker, so each poller gets one full response and then 304s until the next write. That one claim was not a full scan, so it returned nothing and the entries waited for an unrelated write to the organism: on the preview, 106 rsv-a entries reset by the stale clean-up at 19:47 stayed unclaimed for 20+ min. The clean-up and the edit path now forget the organism's claim cursors before and after their transaction commits, so the next claim scans from the start. Open: a claim whose streaming transaction rolls back leaves its entries before the cursor too; that path runs in an Exposed transaction without Spring synchronization and is not covered here. Portable-to-main: with the claim-cursor commits only Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
- silo.persistence: keep each organism's SILO data directory on a PVC instead of an emptyDir, so a restarted pod serves its last import at once rather than waiting for a full re-import. The Deployment then uses the Recreate strategy, so two importers never write the same directory. The importer still re-downloads and re-preprocesses once after a restart: the previous download it compares hashes against lives in the input emptyDir, which it clears on start. - silo.nodeSelector: overrides podScheduling.nodeSelector for SILO and LAPIS. - silo.inputSizeLimit: sizeLimit of the importer's download emptyDir, so a large download evicts the SILO pod instead of putting the node under DiskPressure. - The silo-importer container now honours resources.organismSpecific, like the silo and lapis containers. All default to the previous behaviour. Portable-to-main: yes Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
queryEngine.lapisAlongsideOrganisms limits the LAPIS/SILO deployments, services, configs and ingress to the listed organisms (empty: all enabled organisms, as before). The preview sets it to sars-cov-2, giving a same-data SILO baseline at ~4M entries at lapis-qe-fixes-preview.loculus.org/sars-cov-2/. The website keeps using the backend's engine. - LAPIS 0.8.8 and main's loculus-silo image (SILO 0.14.3, pinned by digest). - SILO data on a 100Gi PVC; SILO and LAPIS pinned to loculus-agent-1, the node with the most free memory. Download emptyDir capped at 150Gi. - SC2 SILO has not run here before, so the requests are estimates (SILO 20Gi, importer 6Gi, LAPIS 1Gi). The limits (56/32/6Gi, 94Gi in total) leave room for the import peak and stay under the node's available memory. - The SILO run timeout is 4 h. The hard refresh is daily, because each one re-downloads every SC2 entry from the backend under test. ETag changes are still imported as they happen. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
With local-path (WaitForFirstConsumer) the claim stays Pending until a pod mounts it. In wave -5 ArgoCD waited for it to become healthy before applying the SILO Deployment in wave 0, so the sync never finished. Portable-to-main: yes, with 116472b Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…wnload At ~4M SARS-CoV-2 entries a full get-released-data download takes ~21 min, and until now neither side logged anything between "Starting to stream" and the end. - backend: log every 100,000 streamed entries and at the end, with entries/s. - silo-importer: log bytes received and MB/s every 30 s during the download, and the total once a 200 response is complete. Portable-to-main: yes Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…rogress logging) The branch's loculus-silo differs from main's only in the importer's download progress logging; SILO stays 0.14.3 (same SILO_VERSION as main). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Level 1 was chosen as "~14% more bytes" on sequence data, which does not hold for SARS-CoV-2. Measured per thread on preview data, level 6 against level 1: - FASTA (3,000 SC2 sequences, 84 MB): 11.2 vs 27.2 MB, 20 vs 163 MB/s. Each ~30 kb sequence fits deflate's 32 kb window, and only level 6's longer match search finds the previous sequence. - details TSV (100k rows, 88 MB): 6.8 vs 9.5 MB, 181 vs 529 MB/s. - aggregated JSON (12 group-by fields, 36 MB): 1.24 vs 1.67 MB, 338 vs 796 MB/s. With level 6 the gzip sizes match LAPIS 0.8.8 (11.1 MB and 6.7 MB for the same requests). The engine stays faster end to end than LAPIS's 6.8 s and 12.0 s. Cached responses are compressed once, so the extra CPU is paid once per entry. Bulk downloads should use zstd, which is smaller and faster than either level. Portable-to-main: with the query engine only Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Level 8's 258-byte nice match length and 1024-entry hash chain find the previous sequence within deflate's 32 kb window; level 6 stops at 128-byte matches and a 128-entry chain. On 3,000 SARS-CoV-2 sequences (84 MB of FASTA), per thread: level 1: 27.2 MB at 163 MB/s level 6: 11.15 MB at 20 MB/s (LAPIS's level) level 8: 1.40 MB at 51 MB/s level 9: 1.39 MB at 46 MB/s Level 8 is 8x smaller and 2.5x faster than level 6 on these sequences. Tables stay at level 6: there level 8 gains only 2-9% at about half the speed. Portable-to-main: with the query engine only Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Don't merge. A preview of fixes from a review of #7431, stacked on
theo-experiment99. It runs Theo's organisms with ingest on, plus a subsampled mpox for the many-genes and long-sequence case.Fixes
Index / API:
lineageDefinitionreturns{}for root nodes, as LAPIS does.minProportionoutside [0, 1] gets a 400 with SILO's message.Error; an OOM exits the JVM (-XX:+ExitOnOutOfMemoryError).Projector / DB:
sequence_entries_lateral_view(V1.38), instead of aggregating all ofexternal_metadatafor every batch.backendDatabasePoolSize.query_changelogseqs are assigned per organism in commit order (V1.39), so replicas can no longer skip a change.Preview config
subsample_fraction: 0.2, about 2.7k of 13.4k sequences.-XX:MaxRAMPercentage=80and GC logging.max_connections=200,shared_buffers=1GB.🤖 Generated with Claude Code
https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
🚀 Preview: https://qe-fixes-preview.loculus.org