feat(scratch): publish a cache to the remote, pick one up cold - #433
Closed
raphaelvigee wants to merge 1 commit into
Closed
raphaelvigee wants to merge 1 commit into
raphaelvigee wants to merge 1 commit into
Conversation
The point of the whole feature in CI, where every runner starts cold. A slot's
contents are published as immutable snapshots under
`scratch/v1/<slot>/<scope>/<gen>.<hash>.tar.gz`, and a cold runner picks up the
newest one for its branch — falling back to the branch it forked from.
heph tool scratch push --all --producer "$GITHUB_RUN_ID"
heph tool scratch pull --all # builds do this on their own
## Pull is automatic, push is a command
A pull is read-only, costs one list plus one fetch, and every way it can fail —
no entry, remote down, corrupt meta, transfer dies halfway — degrades to a cold
build. So a build does it on its own.
A push is none of those. It is expensive, it mutates shared state, and whether a
given job's cache state deserves to become the branch's published head is a
CI-policy question — answered far better by an `if:` condition in a workflow than
by a heuristic inside heph. So it never happens as a side effect of building.
## Why "latest" is not a pointer
The tempting design is a mutable HEAD object naming the newest snapshot. It is
wrong twice over, and the structural reason is the one that matters: **one cache
serves many branches at once**. A remote holds a live lineage for master, one per
open PR, one per long-running branch, all advancing concurrently and all
legitimately different — there is no single "latest" for a pointer to name. A
pointer *per* branch models that and creates an unbounded set of mutable objects
with no way to relate the heads a cross-branch restore has to compare.
The store could not maintain one safely anyway: `RemoteCacheBackend` has
open_read/open_write/exists/list_names — no compare-and-swap and no delete. Even
for a single branch, two jobs finishing together race, and the loser can be the
one that finishes last, overwriting the pointer with older content.
So entries are immutable and ordering is carried in the key. Resolution is a
prefix list and a max: no mutation, no coordination, correct under concurrent
writers, and it extends to many branches by listing more than one prefix. The
generation is zero-padded hex and leads the key, so lexicographic order *is*
generation order and a listing sorts without fetching anything.
## Generations, not timestamps
A publish is `parent + 1` within its lineage. That is deliberately not a clock: a
slow runner that picked up generation 5 an hour ago and finishes now publishes 6,
which correctly loses to a chain that has since reached 12. A timestamp would
have that backwards, and clock skew across runners makes it worse.
A same-generation fork — two runners both publishing `parent + 1` — is expected,
not an error. The tie-break only has to be deterministic so every reader
converges; ordering by `(generation, bytes, key)` makes it a useful choice too,
preferring the fuller cache among two equally-derived ones.
Republishing identical contents is skipped. The archive is deterministic (entries
sorted), so unchanged contents hash identically, and without this a chain would
grow a generation on every no-op CI run.
## Isolation
Writes go to the current scope and never to a fallback, even the one the cache
was seeded from. A PR job picks up master's snapshot and publishes into its own
lineage, leaving master's head exactly where it was — asserted directly, since it
is what makes this safe to enable on untrusted PR CI.
Symlinks are archived as symlinks rather than followed, so a slot that acquired a
link out of the tree does not publish whatever it points at to every machine that
picks the snapshot up. A snapshot records the absolute path it was produced at,
because a cache whose entries embed paths restores fine elsewhere and is then
*inert* — present and useless, which looks exactly like a hit; the mismatch is
logged so it is diagnosable rather than mysterious.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FArWjycMDyWeSfHHtpgtoU
raphaelvigee
force-pushed
the
raphaelvigee/scratch-tooling
branch
from
August 29, 2026 15:34
ebe2173 to
4bd923d
Compare
raphaelvigee
force-pushed
the
raphaelvigee/scratch-remote
branch
from
August 29, 2026 15:34
6552580 to
0a5edac
Compare
raphaelvigee
marked this pull request as ready for review
August 29, 2026 15:34
Member
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The point of the whole feature in CI, where every runner starts cold. A slot's
contents are published as immutable snapshots under
scratch/v1/<slot>/<scope>/<gen>.<hash>.tar.gz, and a cold runner picks up thenewest one for its branch — falling back to the branch it forked from.
Pull is automatic, push is a command
A pull is read-only, costs one list plus one fetch, and every way it can fail —
no entry, remote down, corrupt meta, transfer dies halfway — degrades to a cold
build. So a build does it on its own.
A push is none of those. It is expensive, it mutates shared state, and whether a
given job's cache state deserves to become the branch's published head is a
CI-policy question — answered far better by an
if:condition in a workflow thanby a heuristic inside heph. So it never happens as a side effect of building.
Why "latest" is not a pointer
The tempting design is a mutable HEAD object naming the newest snapshot. It is
wrong twice over, and the structural reason is the one that matters: one cache
serves many branches at once. A remote holds a live lineage for master, one per
open PR, one per long-running branch, all advancing concurrently and all
legitimately different — there is no single "latest" for a pointer to name. A
pointer per branch models that and creates an unbounded set of mutable objects
with no way to relate the heads a cross-branch restore has to compare.
The store could not maintain one safely anyway:
RemoteCacheBackendhasopen_read/open_write/exists/list_names — no compare-and-swap and no delete. Even
for a single branch, two jobs finishing together race, and the loser can be the
one that finishes last, overwriting the pointer with older content.
So entries are immutable and ordering is carried in the key. Resolution is a
prefix list and a max: no mutation, no coordination, correct under concurrent
writers, and it extends to many branches by listing more than one prefix. The
generation is zero-padded hex and leads the key, so lexicographic order is
generation order and a listing sorts without fetching anything.
Generations, not timestamps
A publish is
parent + 1within its lineage. That is deliberately not a clock: aslow runner that picked up generation 5 an hour ago and finishes now publishes 6,
which correctly loses to a chain that has since reached 12. A timestamp would
have that backwards, and clock skew across runners makes it worse.
A same-generation fork — two runners both publishing
parent + 1— is expected,not an error. The tie-break only has to be deterministic so every reader
converges; ordering by
(generation, bytes, key)makes it a useful choice too,preferring the fuller cache among two equally-derived ones.
Republishing identical contents is skipped. The archive is deterministic (entries
sorted), so unchanged contents hash identically, and without this a chain would
grow a generation on every no-op CI run.
Isolation
Writes go to the current scope and never to a fallback, even the one the cache
was seeded from. A PR job picks up master's snapshot and publishes into its own
lineage, leaving master's head exactly where it was — asserted directly, since it
is what makes this safe to enable on untrusted PR CI.
Symlinks are archived as symlinks rather than followed, so a slot that acquired a
link out of the tree does not publish whatever it points at to every machine that
picks the snapshot up. A snapshot records the absolute path it was produced at,
because a cache whose entries embed paths restores fine elsewhere and is then
inert — present and useless, which looks exactly like a hit; the mismatch is
logged so it is diagnosable rather than mysterious.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01FArWjycMDyWeSfHHtpgtoU
Stack created with GitHub Stacks CLI • Give Feedback 💬