Skip to content

Port upstream 0.64.0: Antigravity cost follow-ups - lower bounds, withheld reads, --refresh (stacked on #610) - #706

Draft
Finesssee wants to merge 2 commits into
codex/integrate-reviewed-ports-20260923from
port/micro-0.64.0-antigravity-cost-followups
Draft

Finesssee wants to merge 2 commits into
codex/integrate-reviewed-ports-20260923from
port/micro-0.64.0-antigravity-cost-followups

Conversation

@Finesssee

@Finesssee Finesssee commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Found by the 0.60.4-0.69.0 port gap audit (entry G3, "Antigravity cost follow-ups not in #610"). Stacked on #610 (codex/integrate-reviewed-ports-20260923); #610 at 15f1091 is merged in.

Summary

Antigravity local cost history is never shown as exact when the read was incomplete or contradicted, and unknown-model pricing is refreshed the way upstream 0.64 does it: in the background for routine reads, and blocking for codexbar cost --refresh.

Upstream reference

CodexBar v0.64.0, steipete/CodexBar PR steipete#3757 (commit 7eebfd4): AntigravityLocalReader, CostUsageFetcher (refreshPricingInBackground || !forceRefresh), CLICostCommand, AntigravityPricingRefreshTests, docs/cli.md.

Ported / Deferred

Ported:

  • Retention (a): a later partial or withheld read of the same window and database roots no longer replaces a previously complete history; only a newer complete read does. Source absence is still reported as unavailable.
  • Withheld reads (c): a read cut off by a hard budget (discovery limits, byte budget, deadline) or whose sources contradict each other (same row or response id with different content) is partial with no published totals.
  • Lower bounds (e): a partial read that decoded trustworthy rows publishes its totals as floors. JSON gains tokensAreLowerBound and costIsLowerBound; CLI text says "at least N"; the Usage & Spend row and the share PNG render ≥N tokens next to the existing localized ≥$X known subtotal. A lower bound is never presented as an exact total, and total_usd stays null.
  • Pricing refresh (d):
    • Routine reads never wait on the network. When the scanned history records a model with no known public price, the desktop Usage & Spend read and serve /cost start one bounded models.dev refresh in the background (upstream app and serve), and the next read picks up the prices. Empty, absent or fully priced history starts no download.
    • codexbar cost --provider antigravity --refresh waits for that refresh and rescans (upstream --refresh). The rescan is kept only when it is at least as complete as the first scan.
    • Refresh targets follow the routing a rescan prices from: each unpriced model and the base of its -tiered/-low/-thinking routing variant, grouped by models.dev provider (google for Gemini, openai for GPT, anthropic for Claude). The request carries no account identity or usage data; a failure leaves the models unpriced.
  • LocalTokenHistorySummary and LocalCostEstimate moved to rust/src/spend_contract/local_history.rs so spend_contract.rs stays under 1000 lines. The unpriced model names are local-only (#[serde(skip)]), never part of the wire contract.

Deferred, with Windows design notes:

  • Event-timestamp pricing (b). Windows design note: Win-CodexBar has one bundled price table plus a cached models.dev snapshot, with no dated price tables to choose from, so there is nothing to price "at the event's timestamp" against.
  • "Unstable evidence" detection in (c). Windows design note: upstream retries SQLite reads immutably and compares them. Win reads inside a read-only transaction with no immutable retry, so only contradiction and hard-budget truncation are covered.
  • General catalog refresh when every model is priced. Windows design note: upstream also runs a TTL refreshIfNeeded for fully priced history. Win's models.dev lookups return nothing once the cached catalog is stale, so those models become unpriced and take the targeted refresh above; models priced by the bundled table need no download.
  • Background refresh from a plain codexbar cost. Windows design note: upstream's CLI also starts a background refresh without --refresh, but the process exits right after printing, so Win starts no download there and documents --refresh instead.
  • Transport-injected refresh tests. Windows design note: the pricing path has no injectable models.dev transport or cache root, so upstream's download-against-a-fixture tests are translated as decision-logic tests (which reads start a refresh, which targets it requests, when a rescan replaces the first scan).
  • Per-day dashboard "≥" ledger rendering and Swift-only menu, locale and quota-window strings: they have no Win counterpart.

Validation

At 317d4fe:

  • cargo +1.98.0 fmt --all --check: clean
  • cargo +1.98.0 clippy --workspace --all-targets -- -D warnings: clean
  • cargo +1.98.0 test -p codexbar: 2246 passed, 0 failed, 1 ignored
  • cargo +1.98.0 test -p codexbar-desktop-tauri: 478 passed, 1 failed. The failure is bootstrap_payload_exposes_every_provider_variant, which reads the machine's real settings (Isolate bootstrap payload test from real settings #684, hermetic fix in Make the bootstrap catalog test hermetic (#684) #711); it is unrelated to this change.
  • pnpm install --frozen-lockfile, pnpm run check-locale (875 keys OK), pnpm run lint (clean), pnpm test (401 passed in 66 files), pnpm run build: all ok
  • Tests added: lower-bound and withheld JSON; retention across partial and withheld reads; unpriced-model recording (not serialized); refresh targets per models.dev provider; refresh offered only for named unpriced models; routine reads of absent, empty, priced and unpriced history (upstream "routine local reads do not wait for pricing" and "empty history starts no download"); explicit refresh rescans only when pricing became available and never replaces a complete scan with a partial one; SQLite lower-bound and contradiction cases; the hard-budget JSONL case now withheld; Usage & Spend lower-bound row values; ≥ token formatting.

Affected areas

  • rust/src/providers/antigravity/ (SQLite and tokscale JSONL readers, cost estimate, refresh targeting, history retention)
  • rust/src/spend_contract/ (local history summary and JSON contract)
  • rust/src/cli/cost.rs (--refresh, lower-bound text), rust/src/cli/serve/data.rs (/cost background refresh)
  • rust/src/core/cost_pricing/claude.rs, rust/src/core/claude_routed_pricing.rs (routed models.dev targets, crate-internal)
  • apps/desktop-tauri/src-tauri/src/commands/usage_spend.rs, apps/desktop-tauri/src/lib/usageSpendSharing.ts, apps/desktop-tauri/src/surfaces/settings/tabs/UsageSpendTab.tsx, apps/desktop-tauri/src/types/bridge.ts
  • docs/CLI.md

UI proof

UI proof (browser-use) at 317d4fe: #706 (comment)

  • Complete history (control): the Antigravity Usage & Spend row shows exact totals with no ≥.
  • Lower-bound history: the row shows ≥$3.00 known · ≥600,000 tokens and ≥$6.00 known · ≥1,200,000 tokens.
  • Both runs: dark theme under auto, and no email or account text.
  • Not covered in the UI: the share PNG and the CLI/serve refresh paths. Unit tests cover them.

…held reads, --refresh)

Partial Antigravity history never replaces a previously complete read, contradicted or budget-truncated reads are withheld, and trustworthy partial reads are published as lower bounds (tokensAreLowerBound / costIsLowerBound). Adds cost --provider antigravity --refresh for one bounded pricing refresh; routine reads never download pricing.

Ports upstream steipete/CodexBar PR steipete#3757 (7eebfd4) in part.
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…icing in the background

Resolve the usage_spend.rs, usageSpendSharing, UsageSpendTab, cost.rs,
local_history.rs and local_sqlite_tests.rs conflicts against #610 at
15f1091 (localized known-subtotal label, tray-panel-only main).

Review fixes:
- Routine reads (desktop Usage & Spend, serve /cost) now start one bounded
  models.dev refresh in the background when the history records a model with
  no known public price, as upstream 0.64 does; the next read reprices.
- Pricing refresh targets route each unpriced model and its routing base
  through the same models.dev providers a rescan prices from (google for
  Gemini, openai for GPT, anthropic for Claude) instead of anthropic only.
- docs/CLI.md no longer calls Antigravity cost token-history only.
- Translate the upstream routine-read and explicit-refresh scenarios.
@Finesssee

Copy link
Copy Markdown
Collaborator Author

Lane A review: fixes at 317d4fe

Merged #610 at 15f1091 into the branch (no rebase) and resolved the conflicts in usage_spend.rs, usageSpendSharing.ts / .test.ts, UsageSpendTab.tsx, antigravity/cost.rs, antigravity/local_history.rs and antigravity/local_sqlite_tests.rs toward #610 (tray-panel-only main, localized known-subtotal label). The lower-bound cell keeps the localized UsageSpendKnownSubtotal template.

Defects found and fixed:

  1. False deferral. The PR body deferred the background routine pricing refresh because "Win has no routine network pricing path". Win does have one (refresh_unknown_models_if_needed: bounded, with a 15-minute attempt window). Upstream 0.64 refreshes in the background for the app and serve (refreshPricingInBackground || !forceRefresh). The desktop Usage & Spend read and serve /cost now spawn one bounded refresh when the history records an unpriced model, and the next read reprices. Empty, absent or fully priced history starts nothing.
  2. Wrong refresh target. The refresh asked models.dev provider anthropic for every model, so the Gemini and GPT models Antigravity records could never be priced by a refresh. Targets now follow the routing a rescan prices from (google, openai, anthropic), including the base of -tiered / -low / -thinking variants. The refresh reports success when any group's models became priced, so a shared-catalog download that priced a later group still triggers the rescan.
  3. Stale docs. docs/CLI.md still said Antigravity cost is "token history only". It now covers list-price estimates, lower bounds, withheld reads and when pricing is refreshed.

Tests added: refresh_targets_route_unpriced_models_and_their_routing_base_by_vendor, pricing_refresh_is_offered_only_for_named_unpriced_models, routine_read_offers_background_pricing_only_for_unpriced_history (absent / empty / priced / unpriced SQLite history), explicit_refresh_rescans_only_when_pricing_became_available (no refresh, offline, repriced, partial or vanished rescan, partial first scan).

Validation at 317d4fe:

  • cargo +1.98.0 fmt --all --check: clean
  • cargo +1.98.0 clippy --workspace --all-targets -- -D warnings: clean
  • cargo +1.98.0 test -p codexbar: 2246 passed, 0 failed, 1 ignored
  • cargo +1.98.0 test -p codexbar-desktop-tauri: 478 passed, 1 failed (bootstrap_payload_exposes_every_provider_variant reads the machine's real settings: Isolate bootstrap payload test from real settings #684, hermetic fix in Make the bootstrap catalog test hermetic (#684) #711; unrelated)
  • pnpm: install --frozen-lockfile, check-locale (875 keys OK), lint (clean), test (401 passed in 66 files), build: all ok

Checked with no change needed: locale keys (the ≥ prefix is language-neutral and the subtotal label stays localized), bridge types (sevenDayTokensLowerBound / thirtyDayTokensLowerBound match the camelCase serde fields; unpriced model names are #[serde(skip)]), provider data siloed, no secrets logged (the new paths add no logging), no new dependencies, all touched files under 1000 lines.

Remaining deviations are listed in the PR body with Windows design notes: no dated price tables (b), no immutable-retry "unstable evidence" check (c), no general TTL catalog refresh when every model is priced, no background download from a plain codexbar cost, and the transport-injected upstream tests translated as decision-logic tests. UI proof of the ≥N tokens row follows in a separate comment.

@Finesssee

Copy link
Copy Markdown
Collaborator Author

UI proof (browser-use)

Head: 317d4fe. Built from that commit with pnpm install --frozen-lockfile and pnpm run tauri:build:debug. The only proof-only patch was the throwaway dirs shim in the root Cargo.toml (plus the Cargo.lock change it causes). It was restored right after the build and never committed.

Setup

  • Isolated kit home: CODEXBAR_PROOF_HOME, USERPROFILE, HOME, APPDATA and LOCALAPPDATA all point into the kit.
  • Settings: Antigravity enabled, theme auto.
  • A seeded Antigravity snapshot with no identity fields, so no provider is contacted.
  • Antigravity local SQLite history in the upstream gen_metadata(idx, data) shape. Each database has 5 rows totalling 600,000 tokens and $3.00 at list price. direct.db is 1 day old; routed.db (-Thinking routing suffix) is 20 days old.
  • Every model has a list price, so no background models.dev download starts.
  • Driven over WebView2 CDP on port 9351, owned by the kit exe, with CODEXBAR_PROOF_MODE=settings:usageSpend.
# Scenario Assertion Result
A0 both No email or account text in the DOM (checked before each screenshot) PASS
A1 both Theme auto resolves dark: prefers-color-scheme: dark, body rgb(28, 28, 30) PASS
A2 complete history (control) Antigravity row shows $3.00 · 600,000 tokens, $6.00 · 1,200,000 tokens, source local Antigravity history · API list-price estimate. Exact totals carry no ≥. PASS
A3 lower-bound history: direct.db also holds one undecodable row (the undecodable_row_beside_valid_rows_yields_a_lower_bound case) Antigravity row shows ≥$3.00 known · ≥600,000 tokens, ≥$6.00 known · ≥1,200,000 tokens, source local Antigravity history · known API list-price subtotal. The partial read is shown as a floor, never as an exact total. PASS

Screenshots are kept locally in the proof kit (port-audit\proof\706\shots\). They show fixture data only:

  • 01-usage-spend-complete.png
  • 02-usage-spend-lowerbound.png

Not covered in the UI

  • The share PNG. Sharing downloads a file through WebView2, which could land outside the kit, and its 100 px cells truncate these strings anyway. The ≥ token formatting it reuses is covered by apps/desktop-tauri/src/lib/usageSpendSharing.test.ts.
  • The CLI cost --refresh rescan and the serve /cost background refresh. These are covered by the Rust tests listed in the PR body.

The PR stays a draft for the maintainer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant