Skip to content

fix(grok): preserve local token totals in Usage & Spend - #4093

Closed
steipete wants to merge 2 commits into
mainfrom
triage/20260921-web-sessions-2
Closed

steipete wants to merge 2 commits into
mainfrom
triage/20260921-web-sessions-2

Conversation

@steipete

Copy link
Copy Markdown
Owner

Grok's successful proxy fallback already carries local token history. Usage & Spend lost it later: requesting a wider history window changed a complete 30-day scan into unknown coverage, so the dashboard and share builder discarded valid totals.

Keep the scan's actual coverage in wider views, publish fresh local tokens when a billing outage retains older quota, and align file selection with the declared local calendar days. Quota timestamps retain their original meaning. The app reuses package-scoped snapshot copying and day formatting; production code is 21 insertions / 25 deletions (net -4).

Verification

Synthetic regressions failed before the fixes:

  • Retained-quota publication: 6 failed assertions across the restored/live cases; the missing-quota control passed.
  • Coverage and scanner windows: 3 tests failed with 16 issues, including a successful quota refresh with 85 in-memory tokens producing no dashboard total or share payload.

Final focused run: 220 tests in 19 suites passed. Assertions cover UsageStore refresh, the actual dashboard request boundary, shared token totals, quota timestamps, calendar boundaries, and account/config isolation. The existing synthetic -32601 → proxy fallback regression also passes. Usage JSON intentionally omits live-only costUsage, so the tests inspect memory and dashboard inputs.

mkdir -p .build/web-sessions-tmp
source Scripts/test_environment.sh
TMPDIR="$PWD/.build/web-sessions-tmp/" swift test --jobs 4 -debug-info-format none \
  -Xswiftc -Xfrontend -Xswiftc -emit-macro-expansion-files -Xswiftc -Xfrontend -Xswiftc none \
  --filter 'GrokBillingFailurePublicationTests|DevinSessionImporterTests|DevinUsageFetcherTests|GrokAccountContextTests|GrokLocalSessionScannerTests|GrokTokenSnapshotProjectionTests|GrokAuthTests|GrokFailedBillingWorkTests|StepFun|GrokRemainingResetsFetcherTests|UsageStoreSupplementalUsageTests|SpendDashboardTokenProvenanceTests|ProviderTransportGenerationTests|ProviderArchitectureGatekeeperTests'
make check

The local debug build disables debug artifacts to work around compiler filesystem-rename stalls. make check passed on the integrated branch: 0/2670 files require formatting; SwiftLint reports 0 violations in 2669 files; repository checks pass. One unchanged process-cleanup fixture timed out once; its focused retry and the final full check passed. Independent Codex autoreview found no actionable P0–P2 findings.

The clean main merge leaves the Grok source/test files unchanged from the validated fix. Changelog and provider documentation are updated. Thanks @Chipagosfinest for the detailed report.

Fixes #3716
Refs #3660
Refs #3762

Keep the actual scanned history window when a dashboard requests more history. Publish fresh local tokens even when failed billing retains older quota data, preserving quota timestamps. Align session scanning with its declared calendar-day window and reuse the shared day formatter.

Fixes #3716
@clawsweeper

clawsweeper Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. labels Sep 28, 2026
@clawsweeper

clawsweeper Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codex review: needs maintainer review before merge. Reviewed September 28, 2026, 1:40 AM ET / 05:40 UTC.

ClawSweeper review

What this changes

The branch preserves Grok local token totals in wider Usage & Spend views and during billing failures, aligns session scans with local calendar days, and adds tests and documentation.

Merge readiness

✅ Ready for maintainer review

Keep open. Current main still drops established Grok token totals when the dashboard requests a wider history window. This focused patch addresses that downstream gap, which the previously tested proxy fallback alone did not resolve.

Priority: P2
Reviewed head: a0416726790c5af0eac450bd5b6a433858e8a1af

Review scores

Measure Result What it means
Overall readiness 🐚 platinum hermit (4/6) A small, coherent repair with relevant synthetic regressions; the owner-authored PR has no external-contributor proof gate.
Proof confidence 🌊 off-meta tidepool Not applicable: The OWNER-authored PR is exempt from the external-contributor real-run gate. Reported passing synthetic tests exercise the changed Grok scanner and UsageStore through dashboard and share inputs, but no live app result is supplied. costUsage remains excluded from stored JSON, so no stored-data migration is involved.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Not applicable Not applicable: The OWNER-authored PR is exempt from the external-contributor real-run gate. Reported passing synthetic tests exercise the changed Grok scanner and UsageStore through dashboard and share inputs, but no live app result is supplied. costUsage remains excluded from stored JSON, so no stored-data migration is involved.
Evidence reviewed 10 items Current-main gap: The base revision projects a 30-day Grok scan into any requested wider window and marks its coverage unestablished. The dashboard requests its all-time scan window at the Grok capture boundary.
Dashboard boundary: The dashboard captures Grok's token projection with its wider scanDays request and treats a nil projection as a confirmed empty source.
Introduced projection repair: The introduced guard returns the published snapshot with its actual coverage when the requested window is wider; only narrower windows are projected.
Findings None None.
Security None None.

How this fits together

The Grok provider combines remote quota results with token counts from local session files. UsageStore publishes those counts to Usage & Spend and shared cards while retaining the quota's original timestamp.

flowchart LR
 A[Local Grok sessions] --> B[Calendar day scan]
 B --> C[Token snapshot]
 D[Remote quota result] --> E[UsageStore publication]
 C --> E
 E --> F[Usage & Spend]
 E --> G[Shared cards]
Loading

Before merge

None.

Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Production and test lines production +21/-25; tests +159/-12 The production repair is small and is accompanied by focused regression coverage.

Technical review

Best possible solution:

Use the scan's true coverage in dashboard inputs and keep freshly read local tokens separate from the age of retained quota data.

Do we have a high-confidence way to reproduce the issue?

Yes, from source: a complete 30-day Grok scan is passed to the dashboard's wider request, where current main marks its coverage unestablished. The added synthetic tests exercise that boundary; this review did not execute them or replay the reporter's account.

Is this the best way to solve the issue?

Yes. Preserving actual scan coverage and publishing fresh local tokens within the existing refresh guards is a narrow repair; the patch retains token-only reporting and quota timestamps.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against 579f68406855.

Labels

Label changes:

  • add P2: The patch addresses missing analytics totals for one provider without evidence of lost session files or a blocked core workflow.
  • add rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🌊 off-meta tidepool and patch quality is 🐚 platinum hermit.
  • add status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: The OWNER-authored PR is exempt from the external-contributor real-run gate. Reported passing synthetic tests exercise the changed Grok scanner and UsageStore through dashboard and share inputs, but no live app result is supplied. costUsage remains excluded from stored JSON, so no stored-data migration is involved.

Label justifications:

  • P2: The patch addresses missing analytics totals for one provider without evidence of lost session files or a blocked core workflow.
  • rating: 🐚 platinum hermit: Overall readiness is 🐚 platinum hermit; proof is 🌊 off-meta tidepool and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: The OWNER-authored PR is exempt from the external-contributor real-run gate. Reported passing synthetic tests exercise the changed Grok scanner and UsageStore through dashboard and share inputs, but no live app result is supplied. costUsage remains excluded from stored JSON, so no stored-data migration is involved.

Evidence

What I checked:

Likely related people:

  • steipete: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • Chipagosfinest: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

steipete added a commit that referenced this pull request Sep 28, 2026
…4093)

Grok local token history now reaches Usage & Spend and share output when x.ai billing is unavailable: wider dashboard requests keep the scan's actual 30-day coverage instead of an unknown horizon that the dashboard rejected, fresh local history publishes before retained-quota early returns, and the scanner clips files to its advertised calendar days so aggregates match daily buckets. Fixes #3716.

Thanks @Chipagosfinest!
@steipete

Copy link
Copy Markdown
Owner Author

Landed on main as 03f4b68 via merge train #4096 (one green CI run for the whole train).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

P2 Normal priority bug or improvement with limited blast radius. rating: 🐚 platinum hermit Good normal PR readiness with ordinary maintainer review expected. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Grok token totals never reach Usage & Spend: x.ai/billing returns -32601 on grok CLI 1.0.25 and the local session summary is dropped

1 participant