Skip to content

Port upstream 0.65.0: price Codex Priority turns from trace evidence - #698

Open
Finesssee wants to merge 6 commits into
port/upstream-0.65.0from
port/micro-0.65.0-codex-priority-trace
Open

Finesssee wants to merge 6 commits into
port/upstream-0.65.0from
port/micro-0.65.0-codex-priority-trace

Conversation

@Finesssee

Copy link
Copy Markdown
Collaborator

Summary

Codex local cost now prices Priority (Fast) turns at the Fast rate even when the session model is a plain gpt-5.x name. Codex records each websocket request in <CODEX_HOME>/logs_2.sqlite; a response.create request with service_tier == "priority" marks the turn as Priority. The cost scanner reads that database read-only and reprices matching session rows when day totals are rebuilt. Backend only: corrected numbers show up wherever Codex cost already shows (tray cost card, Settings usageSpend, codexbar cost). No new setting, surface, or locale key.

Approved design

Design note design-codex-priority-trace.md (approved DEFER item, upstream v0.65.0 item #29), implemented as written, backend only. The optional usageSpend "Priority pricing detected" line was skipped as recommended.

  • Read <CODEX_HOME>/logs_2.sqlite (logs(rowid, ts, feedback_log_body), index idx_logs_ts) through the existing open_readonly_sqlite_connection helper (WAL safe, 250 ms busy timeout, never writes).
  • Same query filter, parsers (websocket request: JSON, Submission sub=Submission { rows, websocket event: response.completed), turn-id extraction, and completed-model retention limit (4096) as upstream.
  • Incremental durable cursor in the cost cache (codex_priority_turns_cursor): last rowid, NTFS file identity, coverage window start, SHA-256 anchors over four sampled rows (a rewritten DB rebuilds), pruning of turns whose source rows Codex deleted, and rowid-chunked scanning that honors cancellation and keeps partial progress.
  • Session rows carry a turn_id (from task_started / turn_id); the Priority overlay is applied when day totals are rebuilt and in quota-window slices, and is never baked into persisted pricing evidence.
  • Missing or unreadable DB keeps current Standard behavior (tracing::debug! only, no bodies, no error). Turn evidence is scoped to sessions under the same CODEX_HOME as the database. Only ids, model names, and timestamps are stored; bodies are parsed in memory and never logged.
  • No new dependencies (rusqlite, sha2 already present). Codex cache schema bumped 3 to 4, so older caches rebuild once.

Upstream reference

Ported / Deferred

Ported: everything above, including cursor, anchors, pruning, retention limit, and -priority pricing through existing Fast multipliers (models without a Fast lane stay Standard, as upstream).

Deferred / differences:

  • Optional usageSpend line (skipped per design).
  • Upstream's process-level memo and anchorRowID/anchorDigest single anchor are replaced by the persisted cost-cache cursor with four sampled anchors and a two-of-four quorum; behavior for a rewritten DB is the same (rebuild).
  • This machine's logs_2.sqlite has no priority rows to compare against (schema and turn.id= prefix format were confirmed read-only), so real-data proof is fixture based.

Validation

Toolchain cargo +1.98.0, build slot-2, E-core wrappers.

  • cargo fmt --all: clean.
  • cargo clippy --workspace --all-targets -- -D warnings: pass (both manifests).
  • cargo test -p codexbar priority_trace: 20 passed (parsers, cold/incremental scan, anchors, rewritten DB, pruning, cancel, missing/corrupt DB, retention bound, scope, end-to-end scan pricing).
  • cargo test -p codexbar (full): 2180 passed, 0 failed, 1 ignored.
  • Tauri crate sources untouched; it compiled under workspace clippy.
  • Frontend untouched; vitest not run.

Affected areas

  • Rust backend (rust/src/cost_scanner, core/jsonl_scanner, codex_costs)
  • Docs (docs/PROVIDERS.md Codex cost note)
  • Tauri shell / React / i18n: none

UI proof

Not applicable (backend only; no UI surface changed).

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 2a7a92a8-c94e-46e3-9c9f-50250d96be81

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Thermo-nuclear review

Spec: design-codex-priority-trace.md (upstream v0.65.0 item #29), read against the tag-pinned upstream CostUsageScanner+CodexPriority.swift, CostUsageScanner.swift and CostUsagePricing.swift::codexPriorityCostUSD. Behavior matches the spec: same query filter, parsers, turn-id extraction, retention limit 4096, read-only WAL-safe open, scope by CODEX_HOME, overlay never persisted into pricing evidence. No file crosses 1000 lines (the new resolver is split into priority_trace.rs / parse.rs / tests.rs). No new dependencies.

Findings

  1. Duplicated "row -> priced model" logic (structural). cache_days::row_pricing_model and quota_windows::slice_from_row each carried the same 8-line pricing_mode == "priority" suffix folding plus the overlay lookup, and row_pricing_model returned a (String, bool) whose flag only existed to feed an any() pre-scan. One canonical helper (source_rows::priced_model) deletes both copies and the bool.
  2. turns was redundant persisted state. CodexPriorityTurnsCursor.turns is exactly "latest row of request_sources[turn]", but four call sites (absorb, prune, prune-partial, overlay) had to keep the two maps in sync. Derive it (cursor.turn(id)) and the invariant cannot break; the cache also shrinks.
  3. completed_order was a Vec evicted with remove(0). After the first 4096 non-priority completions every further completion memmoves the whole queue, on the cold-scan hot path. VecDeque::pop_front is the direct structure.
  4. priority_days scanned rows twice to decide whether anything changed. Building both day maps and comparing them is simpler and needs no flag.

Checked, no change needed

  • Debounce path returns cached days without re-resolving the trace; the next non-debounced scan resolves it, so evidence lags by at most the debounce window. Acceptable.
  • Cold scan skip via min(rowid) ... indexed by idx_logs_ts where ts >= window is exact (no monotonic-ts assumption).
  • Long-context Priority rows (>272k input, model without a Fast long-context lane): upstream returns nil cost; here -priority names already flow through the same fallback path, so the overlay inherits existing handling. Pre-existing behavior, left alone.
  • resolve_priority_turns is long (about 110 lines of previous/prune/fresh/fallback states). It is readable and heavily tested; a state enum would move the same concepts around without deleting any, so it stays.
  • Persisted cursor size is not part of the cache budget estimate (only files/days). Bounded by Codex's own log retention plus the 4096 pending cap; left as is.

Fixes for 1-4 are being pushed as "Address thermo review".

@Finesssee

Copy link
Copy Markdown
Collaborator Author

Thermo review fixes landed at d3e857c ("Address thermo review"). Reviewed by Codex gpt-6-luna (xhigh); verified and validated by Claude.

Fixed:

  1. Single source_rows::priced_model helper replaces the duplicated row-to-priced-model logic in cache_days and quota_windows.
  2. Removed the redundant persisted turns map; derived via cursor.turn(id).
  3. completed_order is now a VecDeque with pop_front.
  4. priority_days builds both day maps and compares, no pre-scan.
  5. Coverage floor now advances with the scan window and expired turn evidence is pruned without restarting the scan.
  6. Anchor digest hashes typed SQLite values, so REAL timestamps keep their fractional part (regression test added).
  7. A missing trace database on first scan now marks validation pending so the caller's single debug log fires.

Left: none. UI unchanged (optional UI line stays skipped per the design note).

Commands: cargo +1.98.0 fmt --all --check, cargo +1.98.0 clippy --manifest-path rust/Cargo.toml --all-targets -- -D warnings (clean), cargo test --lib priority (23 pass), cargo test --lib cost (263 pass). Tauri crate and frontend untouched.

Price a Priority day aggregate at the base model's short-context rates
times the Fast multiplier, and price a Priority row outside the Fast lane
at its Standard base like upstream. Keep fork rows as source evidence,
keep a vanished trace database pending once it supplied evidence, bypass
the debounce when the database first appears, read trace JSON leniently
and keep the serialized cursor byte-stable.
@Finesssee

Copy link
Copy Markdown
Collaborator Author

Lane B review: fixes at acc54c4

I reviewed the full diff against port/upstream-0.65.0 and against the tag-pinned upstream v0.65.0 sources (CostUsageScanner+CodexPriority.swift, CostUsageScanner.swift, CodexPriorityDatabasePath.swift) and tests. I then pushed four commits on top of d3e857c8, fast-forward only.

Fixes

  • 14393ef7 Harden Codex Priority trace pricing and cache evidence:
    • Priority day aggregates were priced wrong. A day key such as gpt-5.5-priority sums many Fast requests, so its input often passed 272k. The summed input then either lost the surcharge or switched to long-context rates, which no single Fast request can reach. The new codex_day_aggregate_cost_usd prices a Fast key at the base model's short-context rates for that day, times the Fast multiplier. Astra publishes long-context Fast rates, so it keeps the rule Standard aggregates use. Day totals, the legacy quota-window slices and the summary path share this function.
    • Rows outside the Fast lane kept the surcharge. Upstream charges a Priority request that exceeds the Fast lane (input above 272k, Astra excepted) or has no Fast rate at the Standard cost. The new row_priced_model does the same for source rows and quota-window slices. Cached reads do not count toward the limit.
    • Fork rows lost their evidence. Source-row evidence skips fork-shaped files, so trace evidence never reached forked sessions. Their request rows are now kept in codex_fork_rows, built from the same parsed records as the file's day totals, so trace evidence can price fork-child growth.
    • A vanished database dropped its evidence. Once logs_2.sqlite has supplied evidence and then disappears, the cursor stays pending and keeps its evidence (upstream expectExistingDatabase). The new codex_priority_metadata_key is not advanced while pending, so the cached surcharge survives, as upstream.
    • A newly created database was ignored until the debounce expired. When the metadata key changes to a new sqlite: path, the scan now runs straight away (upstream codexPriorityMetadataKey). Unrelated WAL changes still honour the debounce.
    • Trace JSON was read too strictly. A field of the wrong type now reads as absent instead of failing the whole row. response.completed needs an object response, and the marker must be followed by a JSON object.
    • Text timestamps were mishandled. A non-integer text ts is keyed by its leading YYYY-MM-DD and counts from that local midnight, as upstream. Before, it parsed as nothing and was retained forever.
    • The cursor did not serialize stably. The cursor maps are now BTreeMap, so the serialized cursor is byte-stable and does not defeat the unchanged-payload write skip. Both new cache fields are serde(default), so no further schema bump is needed.
  • e41e424e Tests for the Priority day aggregate and row pricing: 6 in cost_pricing_tests.rs and 3 in source_rows_tests.rs.
  • 0904916c cargo fmt on the cache_days imports.
  • acc54c4f Tests translated from upstream:
    • 24 scanner-level tests in the new priority_trace/tests/scanner.rs (859 lines). They translate CostUsageScannerCodexPriorityTests.swift, CostUsageScannerPriorityTests.swift and the ordinary-path case of CodexPriorityDatabaseIsolationTests.swift. Coverage:
      • gpt-5.4, gpt-5.5 and gpt-5.6 Priority rates, and Astra short and long requests.
      • Completed-response and alias models, and an alias repriced when its completion arrives.
      • Missing database, a model without Fast rates, and long-context rows.
      • Cached-read and cumulative-total limits, and fork-child growth.
      • The debounce bypass and the unrelated-WAL case, and the vanished-database key.
      • Budget pruning that rebuilds Priority days from retained files, and the metadata-appeared truth table.
    • 4 parser tests in priority_trace/tests.rs: fields of another type, a completion without an object response, text timestamps, and a text-timestamp row counted on its local day.

The new tests catch regressions. As a mutation check, I made three temporary changes: turned off the day rebuild in save_codex_cache, updated the metadata key while a resolution was pending, and forced "database appeared" to false. Four tests failed: appearing_database_bypasses_the_scan_debounce, budget_pruning_rebuilds_priority_days_from_retained_files, priority_metadata_appears_only_for_a_new_database and vanished_database_keeps_its_evidence_and_metadata_key. I restored scan.rs afterwards; it is unchanged in the pushed head.

Documented differences from upstream

  • Trace path. The trace database follows CODEX_HOME (else %USERPROFILE%\.codex). This keeps it in the same home as the scanned sessions. Upstream defaultURL always uses the ambient home.
  • Speed buckets. Two kinds of Priority row land in the Standard bucket of by_speed, with totals equal to upstream: rows without a Fast rate (gpt-5.4-nano) and rows over the Fast limit (the gpt-5.6 long row). Without trace metadata, the split still shows a Standard bucket.
  • Unpriced routing rows. A routing-unpriced row gives a total of 0.0 plus Partial completeness, not a nil total.
  • Pre-existing gaps, not introduced here.
    • gpt-5.4 and gpt-5.5 have no long-context tier on Windows, so the upstream long-context case is split: a gpt-5.5 day-totals test and a Sol cost test.
    • A Standard aggregate above 272k still uses long rates for the whole aggregate.
  • Test mechanics. Upstream now offsets are emulated by ageing last_scan_unix_ms. The Sol brief test uses the built-in rates, which equal the models.dev rates.

Validation (Windows, toolchain 1.98.0)

  • cargo fmt --all --check: pass
  • cargo clippy --workspace --all-targets -- -D warnings: pass
  • cargo test -p codexbar: 2219 passed, 0 failed, 1 ignored
  • cargo test -p codexbar-desktop-tauri: 461 passed, 1 failed. The failure is commands::tests::bootstrap_payload_exposes_every_provider_variant (catalog 79 vs active 78). It reads host settings, fails the same way on main, and this PR does not touch it.

UI proof: not applicable. The change is backend cost pricing only, with no new surface, setting or locale key.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant