Conversation
Reuse successful saves by stamp and mutation identity, preserving exact Unicode text while avoiding redundant decoding and encoding. Compact native row keys under schema 4; preserve report output and source-based migration. Refs #3882, #3323, #3247. Verified synthetic red-to-green regressions, 266 focused core tests, the 24k-row refresh benchmark, make check, and independent review through P2.
|
🦞👀 Pull request received. I will update this pull request when review starts. ClawSweeper review completeClawSweeper finished reviewing this revision. The review result is being finalized. |
|
Codex review: needs maintainer review before merge. Reviewed September 28, 2026, 12:27 AM ET / 04:27 UTC. ClawSweeper reviewWhat this changesThe branch compacts Claude and Vertex cost-cache rows, reuses decoded artifacts after successful saves, and adds cache upgrade and refresh tests. Merge readiness✅ Ready for maintainer review Keep this PR open. Current main reuses unchanged Claude cache artifacts, but it still encodes repeated saves and does not retain a freshly saved decoded artifact. This PR addresses that remaining work, and the related resource-use reports remain open. Priority: P2 Review scores
Verification
How this fits togetherCodexBar reads Claude and Vertex transcripts to calculate local cost history. The scanner stores intermediate JSON artifacts and report memos that feed the menu and spend dashboard. flowchart LR
A[Claude and Vertex transcripts] --> B[Cost scanner]
B --> C[Native cache artifact]
C --> D{File stamp and content match?}
D -->|Yes| E[Reuse decoded rows]
D -->|No| F[Decode or rebuild rows]
E --> G[Cost report]
F --> G
Before mergeNone. Agent review detailsSecurityNone. Review metrics
Technical reviewBest possible solution: Keep stamp-guarded reuse and the versioned transcript rebuild, preserve report-memo and public JSON formats, and measure the first upgrade on a large existing history. Do we have a high-confidence way to reproduce the issue? Yes. Current main's save path always reaches JSON encoding and does not seed the artifact memo, so a subsequent changed refresh can decode the saved artifact again. The PR supplies red-to-green counts, though this read-only review did not execute the app. Is this the best way to solve the issue? Yes. The change extends the existing stamped cache and rebuilds older native artifacts from transcripts; focused tests check the stored-format transition and report totals. AGENTS.md: found and applied where relevant. Codex review notes: model internal, reasoning high; reviewed against 579f68406855. LabelsLabel changes:
Label justifications:
EvidenceWhat I checked:
Likely related people:
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
|
Claude/Vertex cost caches keep the decoded value of each successful save under its committed file stamp, skip re-encoding unmodified values, and use compact row keys (a 24k-row history artifact shrinks from 9.3 MB to 6.9 MB). Schema 3 -> 4 rebuilds once from existing transcripts; Unicode text is preserved exactly. Refs #3882 #3247. Thanks @djbclark for the CPU sample that pinned this down!
Changed Claude/Vertex transcripts invalidated the decoded artifact after each successful save, so the next refresh decoded the whole history again. Repeated saves also encoded unchanged values before comparing bytes. This follows #4053 and @djbclark’s 24k-row sample in #3882.
Successful saves now retain their decoded value, guarded by the committed file stamp. An in-memory mutation identifier allows unmodified loaded values to skip encoding while preserving exact Unicode text on edits. Compact row keys reduce the JSON artifact size. The native cache schema advances from 3 to 4, rebuilding older rows from transcripts once when needed; report memos and public report JSON retain their formats. The scan state and bounded memo bookkeeping are simpler, keeping production growth at zero: 83 source lines added and 83 removed.
Synthetic debug measurements used one transcript with 24,000 rows and a 9.34 MB starting cache:
Both runs wrote zero payload bytes and showed flat RSS across 1,000 unchanged refreshes. End-of-idle RSS was 306.6 MiB before and 320.6 MiB after, so this does not claim a memory reduction or resolution of the multi-hour app reports. These are application payload bytes, excluding filesystem metadata/journaling. Changed transcripts still replace complete artifacts.
Validation passed through the standard repository test target, including app-level menu/dashboard and Claude-swap routing. Controlled benchmark measurements used a temporary SwiftPM core harness importing this checkout and the original repository test files, with dependencies pinned to the repository lockfile; the shared Mac initially blocked full compilation in filesystem renames. All runs used
Scripts/test_environment.sh; no real account, cookie import, Keychain read, app relaunch, or release build was used.Test run with 274 tests in 27 suites passed, plus the matching Linux-target quota regression (1 test in 1 suite passed).Test run with 266 tests in 24 suites passed.Test run with 15 tests in 2 suites passed.make check: zero violations in 2,671 files. One process-cleanup fixture run had timing failures on the shared machine; the retry passed without changing those tests.The benchmark is available through the normal repository test target with
CODEXBAR_REFRESH_BENCHMARK=1 swift test --filter CostUsageClaudeRefreshBenchmarkTests.Refs #3882
Refs #3323
Refs #3247