Skip to content

fix(cost): compact and reuse Claude cache artifacts - #4092

Closed
steipete wants to merge 1 commit into
mainfrom
triage/20260921-perf-memory-2
Closed

steipete wants to merge 1 commit into
mainfrom
triage/20260921-perf-memory-2

Conversation

@steipete

Copy link
Copy Markdown
Owner

Changed Claude/Vertex transcripts invalidated the decoded artifact after each successful save, so the next refresh decoded the whole history again. Repeated saves also encoded unchanged values before comparing bytes. This follows #4053 and @djbclark’s 24k-row sample in #3882.

Successful saves now retain their decoded value, guarded by the committed file stamp. An in-memory mutation identifier allows unmodified loaded values to skip encoding while preserving exact Unicode text on edits. Compact row keys reduce the JSON artifact size. The native cache schema advances from 3 to 4, rebuilding older rows from transcripts once when needed; report memos and public report JSON retain their formats. The scan state and bounded memo bookkeeping are simpler, keeping production growth at zero: 83 source lines added and 83 removed.

Synthetic debug measurements used one transcript with 24,000 rows and a 9.34 MB starting cache:

Measurement Before After
Native artifact bytes 9,338,526 6,914,526
CPU seconds across 12 append refreshes 13.374 9.565
Cache + memo payload bytes per append refresh 12,152,086 9,727,430
Full artifact decodes during 12 appends 11 0
Encodes during 8 identical saves 8 0
Process peak RSS 322.6 MiB 320.6 MiB

Both runs wrote zero payload bytes and showed flat RSS across 1,000 unchanged refreshes. End-of-idle RSS was 306.6 MiB before and 320.6 MiB after, so this does not claim a memory reduction or resolution of the multi-hour app reports. These are application payload bytes, excluding filesystem metadata/journaling. Changed transcripts still replace complete artifacts.

Validation passed through the standard repository test target, including app-level menu/dashboard and Claude-swap routing. Controlled benchmark measurements used a temporary SwiftPM core harness importing this checkout and the original repository test files, with dependencies pinned to the repository lockfile; the shared Mac initially blocked full compilation in filesystem renames. All runs used Scripts/test_environment.sh; no real account, cookie import, Keychain read, app relaunch, or release build was used.

source Scripts/test_environment.sh
swift test --force-resolved-versions --jobs 4 -debug-info-format none --no-parallel --filter 'CostUsage.*Claude.*Tests|CostUsageScannerTests|CostUsageScannerCodexPriorityTests|CostUsageScannerBreakdownTests|CostUsageCancellationTests|CostUsageTimestampTests|CostUsageTimestampOrderTests|CostUsageWindowSummaryTests|CostUsageScannerWhitespaceTests|SpendDashboardClaudeCacheRoutingTests|ClaudeSwapRetainedCacheCompatibilityTests'
CODEXBAR_REFRESH_BENCHMARK=1 swift test --package-path .lane-evidence/core-tests --scratch-path .build --jobs 4 -debug-info-format none --no-parallel
make check
  • Standard focused verification: Test run with 274 tests in 27 suites passed, plus the matching Linux-target quota regression (1 test in 1 suite passed).
  • Core harness: Test run with 266 tests in 24 suites passed.
  • Final benchmark/regression run: Test run with 15 tests in 2 suites passed.
  • Red → green: the reuse regression observed 4 decodes/6 encodes before, versus 0/3 after; compact-row assertions also failed before and pass after. A separate Unicode regression caught and fixed a value-equality shortcut before landing.
  • make check: zero violations in 2,671 files. One process-cleanup fixture run had timing failures on the shared machine; the retry passed without changing those tests.
  • Independent Codex review through P2: scoped-clean.

The benchmark is available through the normal repository test target with CODEXBAR_REFRESH_BENCHMARK=1 swift test --filter CostUsageClaudeRefreshBenchmarkTests.

Refs #3882
Refs #3323
Refs #3247

Reuse successful saves by stamp and mutation identity, preserving exact Unicode text while avoiding redundant decoding and encoding. Compact native row keys under schema 4; preserve report output and source-based migration.

Refs #3882, #3323, #3247. Verified synthetic red-to-green regressions, 266 focused core tests, the 24k-row refresh benchmark, make check, and independent review through P2.
@clawsweeper

clawsweeper Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. labels Sep 28, 2026
@clawsweeper

clawsweeper Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codex review: needs maintainer review before merge. Reviewed September 28, 2026, 12:27 AM ET / 04:27 UTC.

ClawSweeper review

What this changes

The branch compacts Claude and Vertex cost-cache rows, reuses decoded artifacts after successful saves, and adds cache upgrade and refresh tests.

Merge readiness

✅ Ready for maintainer review

Keep this PR open. Current main reuses unchanged Claude cache artifacts, but it still encodes repeated saves and does not retain a freshly saved decoded artifact. This PR addresses that remaining work, and the related resource-use reports remain open.

Priority: P2
Reviewed head: 07af35cdc08ad8c60fea8ac9aa74cd5e6046a845

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) The focused cache change has substantial regression and benchmark evidence; the large-history measurements use a synthetic setup.
Proof confidence 🐚 platinum hermit (4/6) Not applicable: The owner-authored PR is exempt from the external-contributor proof gate. Its supplied synthetic measurements exercise the production scanner with transcript files and report after-change decode, encode, and size results; schema 3 rebuild coverage verifies the stored-cache upgrade in a fixture, while no live app upgrade was shown.
Patch quality 🦞 diamond lobster (5/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Not applicable Not applicable: The owner-authored PR is exempt from the external-contributor proof gate. Its supplied synthetic measurements exercise the production scanner with transcript files and report after-change decode, encode, and size results; schema 3 rebuild coverage verifies the stored-cache upgrade in a fixture, while no live app upgrade was shown.
Evidence reviewed 6 items Current main still has the remaining work: The pinned main source uses schema 3, decodes on an artifact-memo miss, and calls the JSON writer on every save without seeding the decoded memo.
Introduced cache behavior: The introduced hunk adds compact row keys, schema 4, a mutation identifier, and stamp-guarded retention of successful saves.
Upgrade and reuse coverage: Introduced tests cover schema 3 rebuilding from transcripts, matching report totals, repeated saves, external replacements, cancellation, and exact Unicode persistence.
Findings None None.
Security None None.

How this fits together

CodexBar reads Claude and Vertex transcripts to calculate local cost history. The scanner stores intermediate JSON artifacts and report memos that feed the menu and spend dashboard.

flowchart LR
  A[Claude and Vertex transcripts] --> B[Cost scanner]
  B --> C[Native cache artifact]
  C --> D{File stamp and content match?}
  D -->|Yes| E[Reuse decoded rows]
  D -->|No| F[Decode or rebuild rows]
  E --> G[Cost report]
  F --> G
Loading

Before merge

None.

Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Code and test delta production +83/-83; tests +252/-19 The cache change has no net production-line growth and adds focused regression and benchmark coverage.

Technical review

Best possible solution:

Keep stamp-guarded reuse and the versioned transcript rebuild, preserve report-memo and public JSON formats, and measure the first upgrade on a large existing history.

Do we have a high-confidence way to reproduce the issue?

Yes. Current main's save path always reaches JSON encoding and does not seed the artifact memo, so a subsequent changed refresh can decode the saved artifact again. The PR supplies red-to-green counts, though this read-only review did not execute the app.

Is this the best way to solve the issue?

Yes. The change extends the existing stamped cache and rebuilds older native artifacts from transcripts; focused tests check the stored-format transition and report totals.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against 579f68406855.

Labels

Label changes:

  • add P2: The PR targets a measured Claude cost-refresh performance path with limited scope; the broader resource-use reports remain separately unresolved.
  • add rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🐚 platinum hermit and patch quality is 🦞 diamond lobster.
  • add status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: The owner-authored PR is exempt from the external-contributor proof gate. Its supplied synthetic measurements exercise the production scanner with transcript files and report after-change decode, encode, and size results; schema 3 rebuild coverage verifies the stored-cache upgrade in a fixture, while no live app upgrade was shown.

Label justifications:

  • P2: The PR targets a measured Claude cost-refresh performance path with limited scope; the broader resource-use reports remain separately unresolved.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🐚 platinum hermit and patch quality is 🦞 diamond lobster.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: The owner-authored PR is exempt from the external-contributor proof gate. Its supplied synthetic measurements exercise the production scanner with transcript files and report after-change decode, encode, and size results; schema 3 rebuild coverage verifies the stored-cache upgrade in a fixture, while no live app upgrade was shown.

Evidence

What I checked:

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • djbclark: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

steipete added a commit that referenced this pull request Sep 28, 2026
Claude/Vertex cost caches keep the decoded value of each successful save under its committed file stamp, skip re-encoding unmodified values, and use compact row keys (a 24k-row history artifact shrinks from 9.3 MB to 6.9 MB). Schema 3 -> 4 rebuilds once from existing transcripts; Unicode text is preserved exactly. Refs #3882 #3247.

Thanks @djbclark for the CPU sample that pinned this down!
@steipete

Copy link
Copy Markdown
Owner Author

Landed on main as daffd77 via merge train #4096 (one green CI run for the whole train).

@steipete steipete closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

P2 Normal priority bug or improvement with limited blast radius. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant