Before submitting
Area
apps/web
Steps to reproduce
- Work in a Claude thread until it holds more than 100K tokens.
- Leave it idle for 70 minutes.
- Come back: "Resume with less context" appears.
Expected behavior
The cheap moment is used, and the offer covers Codex too:
- Before the 1-hour cache expires, idle threads above a size threshold get a compaction offer: as a notification, since the user is usually away, or as an opt-in automatic step in the thread's live process.
- After expiry, the banner also offers "continue in a new thread" with a short brief and a link to the old one; the old thread is archived after the new one's first reply.
- The TTL comes from the API response (
usage.cache_creation.ephemeral_1h_input_tokens / ephemeral_5m_input_tokens; it is 5 minutes on API keys and in overage). For Codex (gpt-6-astra), 30 minutes is treated as a minimum, and Codex threads get the offer as well (T3 can already compact them with compactThread).
- For long agent work, more "Auto-compact after" presets (400K, 600K) and a separate "suggest compaction from N tokens" setting, so a long context stays available when it is needed.
Actual behavior
The banner appears only after 70 idle minutes (Claude Code's own resume threshold, CLAUDE_CODE_RESUME_THRESHOLD_MINUTES), while the cache lives 60 minutes on a Claude subscription and 5 minutes on an API key. By then the cache has expired, so Compact re-reads the whole context at the uncached rate. Codex threads get no offer.
Claude Code's own mechanisms do not cover this. It backs the same 70-minute / 100K dialog with a summary prepared while the cache was warm, and has an idle compaction at 90% of the cache window, but both are server-side experiments: only Anthropic decides which accounts get them, as part of its own testing, and a user has no way to turn them on or influence this. My account (2.1.283) is not enrolled in precomputation (tengu_sepia_moth) or idle compaction (tengu_sunny_locket), but is enrolled in tengu_gleaming_fair_reuse, which shows the dialog only when a precomputed summary exists, so it never appears. Even with the experiment on, the summary is prepared only near the auto-compact threshold.
Measured on Opus 5.5, two idle threads of ~210K tokens, API prices:
| Step |
Cost |
| Continue a cold thread (the first message writes the context to the 1-hour cache at 2x input) |
$1.67, then ~$0.05 per message |
| Compact a cold thread: today's banner, shown after 70 idle minutes while the cache lives 60 |
$0.90 |
| Compact a warm thread |
$0.10 |
| First message after either compaction |
$0.13–0.17, then ~$0.01 |
| New thread with a short brief and a link to the old transcript |
$0.19 for three answers |
Modelled from these numbers, returning to a cold 430K thread for 10 short messages costs $4.40 if continued, $2.12 with today's banner, $0.49 if compacted while warm, and $0.53 via a new thread with a brief and a link. For heavy agent work (~87 requests per message in my session) the auto-compact threshold matters more: five heavy messages cost $56 at 950K, $38 at 600K, $32 at 400K, and $26 at 200K.
Related: #8144 (the banner), #11999 (moves it out of the composer stack), #7590 (cache countdown idea).
Impact
Major degradation or frequent failure
Version or commit
0.0.43-nightly.20260926.2282 (code checked on main @ de251fc)
Environment
Desktop app on Linux (NixOS, Wayland, niri); Claude Code 2.1.283 with Opus 5.5; Codex 0.157.1 with gpt-6-astra
Logs or stack traces
Screenshots, recordings, or supporting files
No response
Workaround
Compact by hand within an hour of the last reply, or lower "Auto-compact after" in the Claude provider settings.
Before submitting
Area
apps/web
Steps to reproduce
Expected behavior
The cheap moment is used, and the offer covers Codex too:
usage.cache_creation.ephemeral_1h_input_tokens/ephemeral_5m_input_tokens; it is 5 minutes on API keys and in overage). For Codex (gpt-6-astra), 30 minutes is treated as a minimum, and Codex threads get the offer as well (T3 can already compact them withcompactThread).Actual behavior
The banner appears only after 70 idle minutes (Claude Code's own resume threshold,
CLAUDE_CODE_RESUME_THRESHOLD_MINUTES), while the cache lives 60 minutes on a Claude subscription and 5 minutes on an API key. By then the cache has expired, so Compact re-reads the whole context at the uncached rate. Codex threads get no offer.Claude Code's own mechanisms do not cover this. It backs the same 70-minute / 100K dialog with a summary prepared while the cache was warm, and has an idle compaction at 90% of the cache window, but both are server-side experiments: only Anthropic decides which accounts get them, as part of its own testing, and a user has no way to turn them on or influence this. My account (2.1.283) is not enrolled in precomputation (
tengu_sepia_moth) or idle compaction (tengu_sunny_locket), but is enrolled intengu_gleaming_fair_reuse, which shows the dialog only when a precomputed summary exists, so it never appears. Even with the experiment on, the summary is prepared only near the auto-compact threshold.Measured on Opus 5.5, two idle threads of ~210K tokens, API prices:
Modelled from these numbers, returning to a cold 430K thread for 10 short messages costs $4.40 if continued, $2.12 with today's banner, $0.49 if compacted while warm, and $0.53 via a new thread with a brief and a link. For heavy agent work (~87 requests per message in my session) the auto-compact threshold matters more: five heavy messages cost $56 at 950K, $38 at 600K, $32 at 400K, and $26 at 200K.
Related: #8144 (the banner), #11999 (moves it out of the composer stack), #7590 (cache countdown idea).
Impact
Major degradation or frequent failure
Version or commit
0.0.43-nightly.20260926.2282 (code checked on
main@ de251fc)Environment
Desktop app on Linux (NixOS, Wayland, niri); Claude Code 2.1.283 with Opus 5.5; Codex 0.157.1 with gpt-6-astra
Logs or stack traces
Screenshots, recordings, or supporting files
No response
Workaround
Compact by hand within an hour of the last reply, or lower "Auto-compact after" in the Claude provider settings.