Skip to content

DON'T MERGE: query engine review fixes preview (stacked on #7431) - #7435

Draft
corneliusroemer-agent wants to merge 167 commits into
theo-experiment99from
qe-fixes-preview
Draft

corneliusroemer-agent wants to merge 167 commits into
theo-experiment99from
qe-fixes-preview

Conversation

@corneliusroemer-agent

@corneliusroemer-agent corneliusroemer-agent commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Don't merge. A preview of fixes from a review of #7431, stacked on theo-experiment99. It runs Theo's organisms with ingest on, plus a subsampled mpox for the many-genes and long-sequence case.

Fixes

Index / API:

  • lineageDefinition returns {} for root nodes, as LAPIS does.
  • minProportion outside [0, 1] gets a 400 with SILO's message.
  • Mutations are counted from Postgres rows only up to 20 ids (was 100). The crossover is around 30–50 ids.
  • The index tailer survives any Error; an OOM exits the JVM (-XX:+ExitOnOutOfMemoryError).
  • Readiness waits for each organism's first index load. Liveness is unaffected.
  • Literal regex searches (the website's free-text search) run as a substring search: 240 → 36 ms at 1M distinct values.
  • A full reload drops the old index first and runs one organism at a time. The minimum heap for a reload falls from about 2.2× to about 1.4× the index. The reloading organism answers 503 until it finishes.

Projector / DB:

  • The projector reads accession batches through a per-row lateral view, sequence_entries_lateral_view (V1.38), instead of aggregating all of external_metadata for every batch.
  • The Hikari pool is sized explicitly at 30 (was the default 10), via Helm backendDatabasePoolSize.
  • A slow reconcile re-marks released accessions dirty at 50/s per organism, with a pass at most every 6 h. It heals rows dropped after errors.
  • query_changelog seqs are assigned per organism in commit order (V1.39), so replicas can no longer skip a change.

Preview config

  • Organisms: the default preview set, plus mpox (Pathoplexus config, preprocessing v28) with subsample_fraction: 0.2, about 2.7k of 13.4k sequences.
  • Backend: 1 replica with 8Gi, -XX:MaxRAMPercentage=80 and GC logging.
  • Dev Postgres: 4Gi limit, max_connections=200, shared_buffers=1GB.
  • Ingest: CPU limit raised to 1 CPU.

🤖 Generated with Claude Code

https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt

🚀 Preview: https://qe-fixes-preview.loculus.org

corneliusroemer-agent and others added 18 commits September 28, 2026 13:25
…nce_entries

The query projector reads each dirty batch with `sequence_entries_view ... where accession in (...)`.
Postgres cannot push an IN list into the view's `GROUP BY` subquery over external_metadata, so every
batch of two or more accessions aggregated the whole external_metadata table (all organisms).

New view sequence_entries_lateral_view (V1.38): same columns and rows, but external metadata is
aggregated per row with LEFT JOIN LATERAL (keeping the GROUP BY, so "no external metadata" stays NULL
rather than jsonb_merge_agg's '{}'). streamReleasedSubmissionsWithCompressedSequences uses it when
an accession filter is given; full streams (get-released-data, full rebuild) stay on
sequence_entries_view, whose plan is better for them. CREATE VIEW only takes AccessShare locks, so
the migration does not wait for running get-released-data streams (a CREATE OR REPLACE of the
existing view would need ACCESS EXCLUSIVE on it).

Measured on PG16 with 1M released entries and 1M external_metadata rows (in cache):
- 2000 random accessions: 250 ms -> 25 ms
- 10 newest accessions (the usual "just approved" batch): 230 ms -> 0.1 ms
- full organism stream: unchanged (still uses sequence_entries_view; through the lateral view it
  would be ~15% slower, same md5 of the output)
The generic plan with bind parameters (as Exposed sends them) is the same nested loop.

SequenceEntriesLateralViewTest compares both views row by row (entries without, with one and with
two external metadata updaters, revised and revoked entries, unprocessed entries) and the
accession-filtered released data against the full stream, so the definitions cannot drift apart
silently.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…s Hikari's default 10)

Nothing configured the pool, so each backend replica ran Hikari's default of 10 connections. The
query engine adds bulk readers to that pool: up to 6 connections per metadata export, 4 per sequence
download, ~9 for the projector while it drains (4 writers, each also allocating ids on a second
connection, plus the leader-lock connection) and 5 per organism while an index loads. On the preview,
8 concurrent metadata exports slowed every DB-backed request 4-5x.

The pool size is now 30 (Helm: backendDatabasePoolSize), with minimum-idle 10 so an idle replica keeps
as many connections open as before. Downloads fetch chunk by chunk and give the connection back
after each chunk, so with 30 connections a waiting request waits for one chunk of another request,
not for a whole download; the 30 s Hikari connectionTimeout is only reachable when most connections
are held for long (e.g. index loads of several organisms at startup).

Budget: replicas x 30 plus other clients must stay below Postgres max_connections (default 100),
so 2 replicas (Pathoplexus) leave ~40 for everything else.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… at a slow rate

A projection can go stale without any later change to its source rows: the one-by-one fallback drops
an accession from the dirty queue after any error (transient ones included), and a trigger can be
missed (e.g. writes with session_replication_role = replica). Nothing recomputed such entries except
a full rebuild.

Each projector run now marks the next released accessions of every organism dirty, in accession
order with a per-organism cursor that wraps around, at loculus.query-engine.reconcile-accessions-per-second
(default 50, 0 turns it off). One pass takes (released accessions / 50) s: ~21 min for Pathoplexus'
largest organism (63k), ~5.5 h at 1M. Unchanged entries cost a read and a projection but no write
(change-only upserts; sequence data skipped by source_hash), and the queue gains at most one batch
per run, so live changes never wait behind a whole-organism sweep. The cursor query is an index range
scan (0.2 ms for 50 accessions at 1M rows).

The drain now logs only runs that changed something, as the reconcile marks accessions every run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ing on submissions

Inserting a dirty-queue key that an uncommitted submission transaction has also inserted waits for
that transaction. A reconcile step that waited while holding its own freshly inserted keys could
form the same deadlock cycle as a group rename during an approval. With lock_timeout = 100 ms (below
deadlock_timeout, 1 s) the step is abandoned before a cycle can be detected, and retried from the
same cursor in the next run.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… hours

At 50 accessions/s per organism a small organism finished its reconcile pass within seconds and
started over at once, so the reconcile re-projected it continuously and never idled. A new pass now
starts at most every loculus.query-engine.reconcile-pass-interval-minutes (default 360) after the
previous one started; large organisms (a pass longer than that, e.g. 1M at 50/s = 5.5 h) run back to
back as before. After a restart the first pass starts immediately.

The cursor query stays cheap with organisms interleaved in accession order (one accession sequence
for all organisms), measured at 1M rows split over 5 organisms of 200k plus one of 10k interleaved
at 1%: 0.1-0.4 ms per step for a 200k organism, 1.6-2 ms for the 1% one, 0.03 ms at the end of a
pass; an organism clustered at the start of the key range gets the organism index (0.2-0.3 ms).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The engine wrote `null` for lineage nodes without parents or aliases. LAPIS
writes `{}`, and the website's lineageDefinition schema (types/lapis.ts)
requires an object per node, so the null roots failed validation and broke
lineage autocomplete.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Out-of-range values were accepted: -1 returned every mutation, 2 returned
nothing. LAPIS (via SILO) answers 400 with "Error from SILO: Invalid
proportion: minProportion must be in interval [0.0, 1.0]"; the engine now
returns the same status and message. NaN is rejected as well.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Mutation and insertion requests over <=100 ids loaded the rows from Postgres
and counted them directly. That costs per id (~0.04 ms locally, 0.2-0.4 ms
measured on the preview's real rows) while the bitmap path is a flat 1-3 ms
at 60k-1M entries, so from roughly 30-50 ids on the "fast path" was the
slower one (preview: +20-40 ms at 100 ids).

In-process probe, synthetic SARS-CoV-2-shaped rows, median ms
(bitmap | rows from local Postgres):

  N=60k  k=1   1.1 | 0.24    k=10  0.9 | 0.55    k=100 1.4 | 3.9
  N=1M   k=1   2.1 | 0.11    k=10  3.1 | 0.46    k=100 2.4 | 3.9

Rows stay much faster for a single sequence (the website's details page),
so the limit drops to 20 instead of removing the path. The complement path
(all but <=100 ids) keeps its own limit of 100.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
tick() caught only Exception. scheduleWithFixedDelay cancels all later runs
of a task that throws and keeps the throwable in a discarded future, so any
Error (StackOverflowError, OutOfMemoryError, ...) froze that organism's
index for good without a log line.

tick() now logs every Throwable at ERROR and keeps ticking, except
OutOfMemoryError, which is logged and rethrown: the backend JVM options get
-XX:+ExitOnOutOfMemoryError so an OOM anywhere (index load, a large request)
restarts the pod instead of leaving it half-working.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…loaded

The pod reported Ready as soon as Spring started, while the indexes were
still building, so every rollout served 503 "initializing" to LAPIS
requests routed to the new pod (18-30 s at 1M entries per the spec).

A queryIndex health indicator is OUT_OF_SERVICE until every organism's
index has finished its first load, and is added to the readiness group
only: liveness and the startup probe are unaffected, so a slow build never
restarts the pod. Later reloads keep serving the old index and do not
touch readiness. The indicator is registered even with the query engine
disabled (then always UP), since a health group must name an existing
contributor.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The website's free-text and substring search sends `field.regex` as
"(?i)" plus the backslash-escaped search text, a new pattern per search
value. Each new pattern is evaluated with RE2J against every distinct value
of the column, under the organism's read lock.

Patterns that are an ASCII literal (optionally case-insensitive, with only
escaped metacharacters) now use a plain substring search with RE2's ASCII
case folding (including U+212A and U+017F); everything else still goes
through RE2J. A randomized test checks agreement with RE2J.

New-pattern cost on an accession-like column, median:
  27k distinct values:  6.7 ms -> 1.1 ms
  1M distinct values:   240 ms -> 36 ms

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A full reload (changelog backlog above max(50k, 10% of the index), e.g.
after a projection rebuild) built the new index while the old one kept
serving, so the heap held two copies of the organism's index plus loader
buffers at the peak. The old index is now dropped first; that organism
answers 503 until the reload finishes (1.5 s for 300k synthetic
SARS-CoV-2-shaped entries here, 18-30 s at 1M per the spec).

Smallest -Xmx at which a reload of a 276 MB index (300k entries) succeeds:
before 620-640 MB, after <= 380 MB.

Readiness now tracks the first load per organism instead of the presence
of an index, since every replica reloads at the same time and gating
readiness on it would take all of them out of service together.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each organism's tracker reloaded on its own thread, so a projection
rebuild across organisms reloaded them all at once: one index rebuild plus
~160k parsed loader rows and 4 pool connections per organism. A
service-wide lock now serialises reloads. The old index is dropped only
once an organism's turn comes, so organisms waiting in the queue keep
serving their current index. First loads at startup stay concurrent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ndJvmExtraOpts

For the PR #7431 query-engine scale test on a Pathoplexus preview: the in-cluster
development Postgres had a hardcoded 2Gi limit, stock settings and no volume, and the
backend JVM flags could not be extended from values. All new keys default to the
previous rendering.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…e query-engine preview

Preview-only values for reviewing PR #7431 with the H/G fixes on a Loculus preview.

- mpox ported from Pathoplexus (schema, preprocessing v28 with nextclade
  nextstrain/mpox/all-clades@2026-07-07--14-07-11Z, reference genomes) plus the
  mpoxOutbreakLineage definition, added to defaultOrganisms so Theo's default
  set stays. Ingest from NCBI taxon 3431483 with subsample_fraction 0.2
  (~2.7k of 13.4k records).
- Backend: 1 replica, 8Gi, MaxRAMPercentage=80 and GC logging via backendJvmExtraOpts.
- Dev Postgres: 4Gi limit, max_connections=200, shared_buffers=1GB, 1Gi /dev/shm.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
Replicas tail query_changelog by seq and skip a gap after 10 s. With a global bigserial, a seq that
commits more than 10 s late, or that is still uncommitted when fullLoad reads max(seq), was never
applied to that replica's index, and the reconcile job does not heal it (an unchanged re-projection
writes no changelog row).

ProjectionWriter now takes the organism's query_engine_state row lock first and then inserts
seq = max(seq of the organism) + row_number in a separate statement, so seqs of an organism become
visible in commit order and without holes. V1.39 makes the primary key (organism, seq); the bigserial
default stays for backends on older code during a rolling update.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…fails

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…ngest fits the 1800 s deadline

Mpox parses the whole ~2.5 GB NCBI package before subsampling; Pathoplexus
production does that within 1800 s on 1 CPU. The request stays at 50m.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
@corneliusroemer-agent corneliusroemer-agent added preview Triggers a deployment to argocd don't merge Experiment or preview only; do not merge labels Sep 28, 2026
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
@claude claude Bot added backend related to the loculus backend component deployment Code changes targetting the deployment infrastructure labels Sep 28, 2026
actions-user and others added 6 commits September 28, 2026 12:17
Pure re-indentation and line unwrapping from `yamlfmt` 0.21.0; the parsed
YAML is identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…sts 3Gi so a rolling update can schedule

A push that changed values recreated the dev Postgres empty while the old backend kept
running against it (relation ... does not exist), and the new 8Gi-request backend pod stayed
Pending next to the old one. The 8Gi limit is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
… backoff

A reload that failed partway left the tracker with no index, and the next tick then ran a plain full load outside
reloadLock. That broke the one-copy-at-a-time invariant the heap sizing relies on, just when something was already
wrong, and retried every tail interval with a stack trace each time.

Every load of an organism after its first now takes reloadLock; first loads at startup stay concurrent. A failed
load waits 1 s before the next attempt, doubling up to 5 min, and the backoff resets on success. Tail failures keep
retrying every tick, so a database blip does not delay release-to-queryable.

Also corrects the stale pool-size note on LOAD_READERS.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
… other error

The rethrow did not do what its comment said. A heap OOM never reaches the catch: -XX:+ExitOnOutOfMemoryError
terminates the JVM at the failed allocation. An OOM thrown by Java code, such as "Required array length is too
large" from collection growth, was caught and rethrown, which silently cancelled the organism's scheduled task
while the JVM kept running: after a dropped index, that organism answered 503 for good with readiness UP.

Such an error allocated nothing, so the heap is intact, and it is usually deterministic, so a restart would not
fix it. Halting the JVM would turn it into a crash loop that takes submission down with a single replica, and
would also kill a test JVM. It now goes through the Throwable branch: logged, and retried with the load backoff.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
Exposed's transaction {} retries every SQLException up to its default of 3 attempts, so a 100 ms lock timeout in
the reconcile marking cost up to 300 ms of the leader tick, each attempt logging Exposed's WARN with a stack trace.
The next projector run already retries the step from the same cursor, so the transaction now makes one attempt.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
corneliusroemer-agent and others added 30 commits September 29, 2026 04:10
…thout a preflight

The browser's conditional LAPIS cache set If-None-Match from JavaScript on every repeat. A header set from JS makes
a cross-origin GET non-simple, so each repeat was preceded by an OPTIONS preflight: a full network round trip, and
the preflight cache is per URL, so every new filter combination paid it.

The browser's axios instance now caches POST only. GETs go straight to the network adapter (XHR), which uses the
browser's HTTP cache: the query engine sends Cache-Control: no-cache with an ETag, so the browser stores the
response and revalidates it with its own If-None-Match, which needs no preflight. XHR surfaces the revalidation's
304 as a 200 with the stored body, and nothing in the website relies on seeing a 304 or the x-loculus-lapis-cache
header for a GET. POSTs are preflighted anyway (JSON content type) and no HTTP cache stores them, so they keep the
JS cache. The server-side render cache is unchanged and still caches GET and POST.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018WBqhtsVLbtxVajax5g1dt
…nly changed records

Every run wrote the whole NCBI corpus to disk several times: seqkit's reformatted
sequences.fasta, then sequences.ndjson, then the metadata TSVs and ndjson, and
prepare_files read the full sequences.ndjson again. For SARS-CoV-2 that is ~85 GB
written twice and read about four times, whether or not anything changed.

For non-segmented organisms, select_changed_records replaces format_ncbi_metadata,
filter_out_depositions, calculate_sequence_hashes, prepare_metadata, metadata_filter
and compare_hashes. It reads genomic.fna once to hash each sequence and note where its
record is, streams the data report in order through the same per-record functions the
old scripts used (now factored out of them), and writes the compare_hashes outputs plus
the metadata and sequences of the records to submit or revise. prepare_files runs
unchanged on those. On 100k SARS-CoV-2 records with a routine mix of changes, every
backend-facing output is byte-identical to the old pipeline, and the rules write 94 MB
instead of 6.2 GB. get_previous_submissions no longer waits for the prepared metadata,
so it runs alongside the download.

Segmented organisms keep the old rules: grouping recomputes the hash of a group from
all its segments, so records can't be dropped before grouping.

construct_submitted_dict now streams the previous submissions and keeps two versions
per Loculus accession instead of every entry (~1 kB per previous submission, GBs for
SARS-CoV-2), with the same result, including which version wins a tie and the curated
flag of revoked accessions. Both paths use it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…cords

With mirror_format: zst, the fetch rule decompressed the mirror's genomic.fna.zst to disk (after
seqkit grep for released_after) only for select_changed_records to read it once and seek back into
it for the changed records. For SARS-CoV-2 that is an ~85 GB write and read per run, for a file
that is 2.8 GB compressed.

Non-segmented organisms without a minimizer now keep the download compressed and
select_changed_records decompresses it itself (stdlib compression.zstd, ~6 GB/s here): once to
hash, once more to extract the changed records by byte range, since seeking forward in the
decompressed stream is as fast as reading it. The release filter moves from seqkit grep into the
parser (--fasta-ids = released_accessions.txt): records outside it are skipped before their
sequence is joined or decoded, so an invalid header outside the filter still does not fail the run.
The filter ids are the index's keys, so they cost no extra memory.

Segmented organisms, a configured minimizer and the other fetch paths still write genomic.fna,
which nextclade and calculate_sequence_hashes read.

Byte-identical to 98f8e18 on 100k SARS-CoV-2 records with the fetch rule run from a local
mirror: routine, routine with released_after (31% kept), 50k bootstrap, and subsample + filter.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
bytes.find(b"\n>") costs ~0.7 s/GB on the wrapped NCBI FASTA, because every line starts a partial
match; finding ">" is a memchr at ~0.02 s/GB. The difference was most of the per-GB cost of the
records the release filter skips (for SARS-CoV-2, ~190 of the mirror's ~270 GB) and a quarter of
select_changed_records on the records it keeps.

100k SARS-CoV-2 records: select_changed_records 10.1 -> 8.0 s (routine), 5.5 -> 3.6 s (routine with
31% released after the cut-off). Outputs byte-identical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…ntry

get_previous_submissions peaked at 6.5 GB for 2.84M SARS-CoV-2 submissions: get_submitted read
the whole /get-submitted-metadata body (requests without stream), held every entry as a dict
(~2 kB each with the body), then reopened the output file once per entry to append it.

The response is now streamed (stream=True, 1 MB chunks) into a temporary file and counted as it
goes; as before, it is retried until the count equals x-total-records. The statuses are still
fetched after the entries, so that a version submitted in between has its status rather than
UNKNOWN; the entries are then rewritten from the temporary file with their status through one
file handle. That costs one extra write and read of the file (~0.5 GB at 2.84M), not memory.

The output is byte-identical (mock backend at 100k and 50k entries, with an interrupted first
stream, and with none). Peak RSS at 100k: 254 -> 168 MB; per entry 2.2 -> 1.3 kB. What is left is
get_sequence_status parsing the whole /get-sequences body at once (~1.3 kB per entry transiently,
~380 B kept), which streaming this endpoint does not touch.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
With get_submitted streamed, get_sequence_status was the peak of get_previous_submissions:
response.json() built the list of every sequence entry (~1.3 kB each while parsing, ~3.7 GB at
2.84M SARS-CoV-2 entries) before the {accession: {version: status}} map was filled from it.

An object_hook now records each entry's status as the parser builds the entry and drops the
entry, and the map is keyed by "accession.version" with interned statuses (93 instead of 377 B per
entry). get_submitted is its only caller. On a decoding error the result is empty, as before.

Mock backend, previous_submissions.ndjson byte-identical at 100k and 50k, after a retried stream
and with no entries. Peak RSS of get_submitted at 100k: 168 -> 90 MB; slope 1.3 -> 0.6 kB per
entry, ~1.7 GB extrapolated to 2.84M (6.3 GB before streaming). Also close a 423 response before
retrying, now that GETs may stream.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
… preview (~3.83M)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The ingest calls /get-sequences only to learn the status of every previous
submission, which it then joins onto the get-submitted-metadata entries. At
2.84M SARS-CoV-2 entries that second call is the ingest's memory peak and
over a minute of wall time. With the status in the same response, the ingest
can drop it.

The status is the value of sequence_entries_view.status, but computed with a
correlated subquery that reads sequence_entries_preprocessed_data only for
entries that are neither released nor revocations. Selecting the view's
status column instead makes Postgres hash-join the whole preprocessed-data
table: on 1M synthetic entries that is +1.3-1.7 s and 145 MB of temp files
per call, against +0.06 s for the subquery (4.43 s baseline). A test checks
that every status matches what /get-sequences returns.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…ed-metadata

The backend now returns the status with each entry, so get_submitted no
longer calls /get-sequences to build an accession.version -> status map.
That call parsed one JSON document of every entry and was the ingest's
memory peak (2.40 GB live at 2.84M SARS-CoV-2 entries). The entries are
written to disk as they stream in and moved into place, without the second
rewrite pass that added the statuses.

The status now comes from the same snapshot as the metadata. Before, the
statuses were fetched after the entries, so that a version submitted in
between got its status rather than UNKNOWN; that race is gone.

status is written as the last key of each entry, where the files always had
it, so previous_submissions.ndjson is byte-identical to before (checked
against a mock backend at 100k and 50k entries, after a truncated stream,
and for 0 entries). An entry without status (a backend older than this
ingest) fails the run instead of being treated as unapproved.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
A routine SARS-CoV-2 ingest on the preview failed in get_previous_submissions
with 502 Bad Gateway while a backend pod shut down during a rollout, and that
failed the whole run. get_submitted_metadata now fetches the whole stream
again, into a fresh file, after a connection error, a timeout, a 5xx, a
stream cut mid-line or mid-chunk, a count that differs from x-total-records,
or an entry without status (an old backend pod still serving during the
rollout that brings the status field): up to 8 attempts, with 10 s doubling
to at most 120 s between them (~8.5 min of backoff), then it fails. Each
attempt can additionally wait on get_jwt's Keycloak retry (300 s) or the
request timeout. Client errors (4xx) still fail at once. make_request fetches
a new token for every attempt.

A cut stream used to fail the run on the invalid last line, and a short
count was retried every 60 s without end; both now take the same bounded
backoff.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
autoapprove calls approve-processed-data every ~70 s per organism. It
selected PROCESSED entries from sequence_entries_view, whose status is a
CASE over the join with sequence_entries_preprocessed_data, so Postgres
joined the organism's whole table on every call: 6.9 s and ~67 MB of temp
files per call for 3.8M SARS-CoV-2 entries with nothing to approve (61 % of
the preview DB's temp writes).

Every PROCESSED entry is unreleased, so filtering on released_at IS NULL
changes no result. With a partial index on unreleased entries the plan
starts from the backlog: 1,390 -> 12 ms on 750k entries locally. Without
the index Postgres misestimates the predicate (438k rows expected, 1,206
actual) and still scans the preprocessed data.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The default 200 held autovacuum to 15-34 MB/s on the preview's NVMe; one
TOAST vacuum took 352 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The released_at IS NULL predicate on the view (5d7a358) is planned from
table statistics. After a large ingest they are stale: at 3.8M SARS-CoV-2
entries with 1,206 unreleased, Postgres expected 416k and still hash-joined
all of sequence_entries_preprocessed_data (2.4 s per autoapprove call).

Without an accession filter, approve now reads the organism's unreleased
entries from sequence_entries_unreleased_idx first and restricts the view
query to those pairs, bound as two unnest arrays. That is planned as a
nested loop of primary-key lookups regardless of statistics: 4 ms to plan
and 32 ms to run on the same data (an IN list of row literals took 0.4 s
to plan for 1,200 pairs).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
The preview gets all 16 PPX organisms (pathoplexus d58747e, incl. chikungunya) with PPX's own ingest, preprocessing,
EMBL annotation files and ENA deposition config, so the query engine, preprocessing and ingest see a PPX-shaped
deployment next to SARS-CoV-2. sars-cov-2, cchf-multi-ref and enteroviruses are copied resolved from the chart's
defaultOrganisms and render identically (helm template diff: only the values-hash annotations and the ingest
schedules, which are spread over the organism count, change). The pure test organisms (dummy-organism,
dummy-organism-with-files, not-aligned-organism) are dropped; they hold no data.

Only each PPX organism's newest pipeline version is kept: PPX is mid-reprocess into it, and a fresh organism would
otherwise be processed twice. mpox, cchf, west-nile and ebola-sudan move from the upstream config to PPX's and are
reprocessed (~26k entries; mpox gains EMBL files).

Memory requests for the PPX preprocessing and ingest pods are set from their peak working set on PPX staging during
its 2026-09-28 reprocess, so the scheduler sees the bootstrap's real footprint.

Generated by investigations/2026-09-28-loculus-pr7431-query-engine/59-data/port_ppx_organisms.py.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…ine, with small pods

Follow-up to f6fc913ad:
- No new EMBL annotation files anywhere (create_embl_file: false for every organism). The annotations category stays
  only where entries already carry files (cchf, west-nile, ebola-sudan and the unchanged chart organisms), so their
  links still resolve.
- mpox, cchf, west-nile and ebola-sudan keep the pipeline versions they run on the preview (28, 1, 1, 1+2) with PPX's
  config, so none of them is reprocessed. Their reference sequences and genes are identical to PPX's; cchf's
  PPX-only lineage_S stays empty on existing entries until a later bump.
- One preprocessing pod per PPX organism; SARS-CoV-2 preprocessing goes from 20 pods to 2 (config unchanged, the old
  count is in a comment next to it).
- PPX organisms' preprocessing and ingest requests are near typical use (128Mi; mpox 256Mi / 1Gi ingest, rsv 512Mi
  ingest) and limits are the measured peak x 1.3, at least 512Mi (peaks from PPX staging/production and this preview).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
A revision adds a new version of an existing accession, which sorts before
the claim cursor, so only the next full scan (every 10 s) found it. The
revise endpoint now remembers the new versions after its transaction
commits, and the organism's next claim takes them first. The memory is per
backend replica and bounded; anything it misses is still found by the full
scan, which can then run less often.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
… more often

With the default scale factors (vacuum 20 %, analyze 10 %) sequence_entries
went 12 h without an analyze after a 1M-entry ingest, and the stale
statistics chose a 2.4 s plan over a 30 ms one for the approve query.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
… s on the preview

New entries and revisions submitted to the same replica are claimed without
a full scan, which is left to catch stale resets, rolled-back claims and
revisions on other replicas. At 3.8M SARS-CoV-2 entries each full scan takes
3.6-3.9 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Nine of the Pathoplexus organism images exist only in the Pathoplexus
website, so the preview's landing page and menu showed broken images.
Chikungunya is on staging only, so its image comes from the pathoplexus
repo at a fixed commit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…PPX main peak x1.3

Dengue preprocessing was OOM-killed at 768Mi; PPX main's 7-day per-pod peak
is 1,118 MiB. Mpox's 3200Mi left 13% over its 2,834 MiB peak.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
A CPU limit only stops a container from using idle cores: under contention
the request already sets its share. On the preview, ingest was throttled in
27-85% of CFS periods (cchf, rsv, mpox), stretching runs and holding their
memory longer; init containers (config-processor, keycloak theme, flyway)
were capped at 500m. Memory limits are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Both nextclade runs were hard-coded to --jobs=1, so a pod used one core for
alignment however many it had; the only way to scale was more replicas, each
with its own Python process and loaded dataset. nextclade parallelises over
the sequences of a batch, so nextclade_jobs: N uses N cores per pod.

The default stays 1 (no behaviour change). The preview sets 4 for sars-cov-2,
dengue and mpox, which have no CPU limit since 47c10e8.

Portable-to-main: yes (config.py + nextclade.py; values change is preview-only)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…aiting for another write

Deleting preprocessed rows makes their entries claimable again, but they sort
before the claim cursor. The delete bumps the update tracker, so each poller
gets one full response and then 304s until the next write. That one claim was
not a full scan, so it returned nothing and the entries waited for an
unrelated write to the organism: on the preview, 106 rsv-a entries reset by
the stale clean-up at 19:47 stayed unclaimed for 20+ min.

The clean-up and the edit path now forget the organism's claim cursors before
and after their transaction commits, so the next claim scans from the start.

Open: a claim whose streaming transaction rolls back leaves its entries
before the cursor too; that path runs in an Exposed transaction without
Spring synchronization and is not covered here.

Portable-to-main: with the claim-cursor commits only

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
- silo.persistence: keep each organism's SILO data directory on a PVC
  instead of an emptyDir, so a restarted pod serves its last import at once
  rather than waiting for a full re-import. The Deployment then uses the
  Recreate strategy, so two importers never write the same directory. The
  importer still re-downloads and re-preprocesses once after a restart: the
  previous download it compares hashes against lives in the input emptyDir,
  which it clears on start.
- silo.nodeSelector: overrides podScheduling.nodeSelector for SILO and LAPIS.
- silo.inputSizeLimit: sizeLimit of the importer's download emptyDir, so a
  large download evicts the SILO pod instead of putting the node under
  DiskPressure.
- The silo-importer container now honours resources.organismSpecific, like
  the silo and lapis containers.

All default to the previous behaviour.

Portable-to-main: yes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
queryEngine.lapisAlongsideOrganisms limits the LAPIS/SILO deployments,
services, configs and ingress to the listed organisms (empty: all enabled
organisms, as before). The preview sets it to sars-cov-2, giving a same-data
SILO baseline at ~4M entries at lapis-qe-fixes-preview.loculus.org/sars-cov-2/.
The website keeps using the backend's engine.

- LAPIS 0.8.8 and main's loculus-silo image (SILO 0.14.3, pinned by digest).
- SILO data on a 100Gi PVC; SILO and LAPIS pinned to loculus-agent-1, the
  node with the most free memory. Download emptyDir capped at 150Gi.
- SC2 SILO has not run here before, so the requests are estimates (SILO 20Gi,
  importer 6Gi, LAPIS 1Gi). The limits (56/32/6Gi, 94Gi in total) leave room
  for the import peak and stay under the node's available memory.
- The SILO run timeout is 4 h. The hard refresh is daily, because each one
  re-downloads every SC2 entry from the backend under test. ETag changes are
  still imported as they happen.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
With local-path (WaitForFirstConsumer) the claim stays Pending until a pod
mounts it. In wave -5 ArgoCD waited for it to become healthy before applying
the SILO Deployment in wave 0, so the sync never finished.

Portable-to-main: yes, with 116472b

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…wnload

At ~4M SARS-CoV-2 entries a full get-released-data download takes ~21 min, and
until now neither side logged anything between "Starting to stream" and the end.

- backend: log every 100,000 streamed entries and at the end, with entries/s.
- silo-importer: log bytes received and MB/s every 30 s during the download,
  and the total once a 200 response is complete.

Portable-to-main: yes

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
…rogress logging)

The branch's loculus-silo differs from main's only in the importer's download
progress logging; SILO stays 0.14.3 (same SILO_VERSION as main).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Level 1 was chosen as "~14% more bytes" on sequence data, which does not hold
for SARS-CoV-2. Measured per thread on preview data, level 6 against level 1:

- FASTA (3,000 SC2 sequences, 84 MB): 11.2 vs 27.2 MB, 20 vs 163 MB/s. Each
  ~30 kb sequence fits deflate's 32 kb window, and only level 6's longer match
  search finds the previous sequence.
- details TSV (100k rows, 88 MB): 6.8 vs 9.5 MB, 181 vs 529 MB/s.
- aggregated JSON (12 group-by fields, 36 MB): 1.24 vs 1.67 MB, 338 vs 796 MB/s.

With level 6 the gzip sizes match LAPIS 0.8.8 (11.1 MB and 6.7 MB for the same
requests). The engine stays faster end to end than LAPIS's 6.8 s and 12.0 s.
Cached responses are compressed once, so the extra CPU is paid once per entry.
Bulk downloads should use zstd, which is smaller and faster than either level.

Portable-to-main: with the query engine only

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky
Level 8's 258-byte nice match length and 1024-entry hash chain find the
previous sequence within deflate's 32 kb window; level 6 stops at 128-byte
matches and a 128-entry chain. On 3,000 SARS-CoV-2 sequences (84 MB of FASTA),
per thread:

  level 1: 27.2 MB at 163 MB/s
  level 6: 11.15 MB at 20 MB/s (LAPIS's level)
  level 8:  1.40 MB at 51 MB/s
  level 9:  1.39 MB at 46 MB/s

Level 8 is 8x smaller and 2.5x faster than level 6 on these sequences. Tables
stay at level 6: there level 8 gains only 2-9% at about half the speed.

Portable-to-main: with the query engine only

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ToXFvEHcKsAssvd5StXxky

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

backend related to the loculus backend component deployment Code changes targetting the deployment infrastructure don't merge Experiment or preview only; do not merge preview Triggers a deployment to argocd update_db_schema

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants