Skip to content

feat(scratch): publish a cache to the remote, pick one up cold - #433

Closed
raphaelvigee wants to merge 1 commit into
raphaelvigee/scratch-toolingfrom
raphaelvigee/scratch-remote
Closed

raphaelvigee wants to merge 1 commit into
raphaelvigee/scratch-toolingfrom
raphaelvigee/scratch-remote

Conversation

@raphaelvigee

Copy link
Copy Markdown
Member

The point of the whole feature in CI, where every runner starts cold. A slot's
contents are published as immutable snapshots under
scratch/v1/<slot>/<scope>/<gen>.<hash>.tar.gz, and a cold runner picks up the
newest one for its branch — falling back to the branch it forked from.

heph tool scratch push --all --producer "$GITHUB_RUN_ID"
heph tool scratch pull --all      # builds do this on their own

Pull is automatic, push is a command

A pull is read-only, costs one list plus one fetch, and every way it can fail —
no entry, remote down, corrupt meta, transfer dies halfway — degrades to a cold
build. So a build does it on its own.

A push is none of those. It is expensive, it mutates shared state, and whether a
given job's cache state deserves to become the branch's published head is a
CI-policy question — answered far better by an if: condition in a workflow than
by a heuristic inside heph. So it never happens as a side effect of building.

Why "latest" is not a pointer

The tempting design is a mutable HEAD object naming the newest snapshot. It is
wrong twice over, and the structural reason is the one that matters: one cache
serves many branches at once
. A remote holds a live lineage for master, one per
open PR, one per long-running branch, all advancing concurrently and all
legitimately different — there is no single "latest" for a pointer to name. A
pointer per branch models that and creates an unbounded set of mutable objects
with no way to relate the heads a cross-branch restore has to compare.

The store could not maintain one safely anyway: RemoteCacheBackend has
open_read/open_write/exists/list_names — no compare-and-swap and no delete. Even
for a single branch, two jobs finishing together race, and the loser can be the
one that finishes last, overwriting the pointer with older content.

So entries are immutable and ordering is carried in the key. Resolution is a
prefix list and a max: no mutation, no coordination, correct under concurrent
writers, and it extends to many branches by listing more than one prefix. The
generation is zero-padded hex and leads the key, so lexicographic order is
generation order and a listing sorts without fetching anything.

Generations, not timestamps

A publish is parent + 1 within its lineage. That is deliberately not a clock: a
slow runner that picked up generation 5 an hour ago and finishes now publishes 6,
which correctly loses to a chain that has since reached 12. A timestamp would
have that backwards, and clock skew across runners makes it worse.

A same-generation fork — two runners both publishing parent + 1 — is expected,
not an error. The tie-break only has to be deterministic so every reader
converges; ordering by (generation, bytes, key) makes it a useful choice too,
preferring the fuller cache among two equally-derived ones.

Republishing identical contents is skipped. The archive is deterministic (entries
sorted), so unchanged contents hash identically, and without this a chain would
grow a generation on every no-op CI run.

Isolation

Writes go to the current scope and never to a fallback, even the one the cache
was seeded from. A PR job picks up master's snapshot and publishes into its own
lineage, leaving master's head exactly where it was — asserted directly, since it
is what makes this safe to enable on untrusted PR CI.

Symlinks are archived as symlinks rather than followed, so a slot that acquired a
link out of the tree does not publish whatever it points at to every machine that
picks the snapshot up. A snapshot records the absolute path it was produced at,
because a cache whose entries embed paths restores fine elsewhere and is then
inert — present and useless, which looks exactly like a hit; the mismatch is
logged so it is diagnosable rather than mysterious.

Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Claude-Session: https://claude.ai/code/session_01FArWjycMDyWeSfHHtpgtoU


Stack created with GitHub Stacks CLI • Give Feedback 💬

The point of the whole feature in CI, where every runner starts cold. A slot's
contents are published as immutable snapshots under
`scratch/v1/<slot>/<scope>/<gen>.<hash>.tar.gz`, and a cold runner picks up the
newest one for its branch — falling back to the branch it forked from.

    heph tool scratch push --all --producer "$GITHUB_RUN_ID"
    heph tool scratch pull --all      # builds do this on their own

## Pull is automatic, push is a command

A pull is read-only, costs one list plus one fetch, and every way it can fail —
no entry, remote down, corrupt meta, transfer dies halfway — degrades to a cold
build. So a build does it on its own.

A push is none of those. It is expensive, it mutates shared state, and whether a
given job's cache state deserves to become the branch's published head is a
CI-policy question — answered far better by an `if:` condition in a workflow than
by a heuristic inside heph. So it never happens as a side effect of building.

## Why "latest" is not a pointer

The tempting design is a mutable HEAD object naming the newest snapshot. It is
wrong twice over, and the structural reason is the one that matters: **one cache
serves many branches at once**. A remote holds a live lineage for master, one per
open PR, one per long-running branch, all advancing concurrently and all
legitimately different — there is no single "latest" for a pointer to name. A
pointer *per* branch models that and creates an unbounded set of mutable objects
with no way to relate the heads a cross-branch restore has to compare.

The store could not maintain one safely anyway: `RemoteCacheBackend` has
open_read/open_write/exists/list_names — no compare-and-swap and no delete. Even
for a single branch, two jobs finishing together race, and the loser can be the
one that finishes last, overwriting the pointer with older content.

So entries are immutable and ordering is carried in the key. Resolution is a
prefix list and a max: no mutation, no coordination, correct under concurrent
writers, and it extends to many branches by listing more than one prefix. The
generation is zero-padded hex and leads the key, so lexicographic order *is*
generation order and a listing sorts without fetching anything.

## Generations, not timestamps

A publish is `parent + 1` within its lineage. That is deliberately not a clock: a
slow runner that picked up generation 5 an hour ago and finishes now publishes 6,
which correctly loses to a chain that has since reached 12. A timestamp would
have that backwards, and clock skew across runners makes it worse.

A same-generation fork — two runners both publishing `parent + 1` — is expected,
not an error. The tie-break only has to be deterministic so every reader
converges; ordering by `(generation, bytes, key)` makes it a useful choice too,
preferring the fuller cache among two equally-derived ones.

Republishing identical contents is skipped. The archive is deterministic (entries
sorted), so unchanged contents hash identically, and without this a chain would
grow a generation on every no-op CI run.

## Isolation

Writes go to the current scope and never to a fallback, even the one the cache
was seeded from. A PR job picks up master's snapshot and publishes into its own
lineage, leaving master's head exactly where it was — asserted directly, since it
is what makes this safe to enable on untrusted PR CI.

Symlinks are archived as symlinks rather than followed, so a slot that acquired a
link out of the tree does not publish whatever it points at to every machine that
picks the snapshot up. A snapshot records the absolute path it was produced at,
because a cache whose entries embed paths restores fine elsewhere and is then
*inert* — present and useless, which looks exactly like a hit; the mismatch is
logged so it is diagnosable rather than mysterious.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FArWjycMDyWeSfHHtpgtoU
@raphaelvigee
raphaelvigee force-pushed the raphaelvigee/scratch-tooling branch from ebe2173 to 4bd923d Compare August 29, 2026 15:34
@raphaelvigee
raphaelvigee force-pushed the raphaelvigee/scratch-remote branch from 6552580 to 0a5edac Compare August 29, 2026 15:34
@raphaelvigee
raphaelvigee marked this pull request as ready for review August 29, 2026 15:34
@raphaelvigee

Copy link
Copy Markdown
Member Author

Folded into #403 (declare/reference/mount) and #431 (lineages, store management, remote) — the split was finer than the review needed.

@raphaelvigee
raphaelvigee deleted the branch raphaelvigee/scratch-tooling August 29, 2026 15:51
@raphaelvigee
raphaelvigee deleted the raphaelvigee/scratch-remote branch August 29, 2026 15:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant