Skip to content

v0.10.1 integration: wave/0.10.1-next - #6782

Draft
Hmbown wants to merge 261 commits into
mainfrom
wave/0.10.1-next
Draft

Hmbown wants to merge 261 commits into
mainfrom
wave/0.10.1-next

Conversation

@Hmbown

@Hmbown Hmbown commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Current head: ce1ecc8dc28c7c634909a0fad35d1e6a680f9d48 (tree bdbd41bd1864)

A fast-forward from b42e145e4. It adds five locally qualified batches (3–6). Receipts are in opus55-0101-completion-20260930/batch{3,4,5,5b,6}/.

  • Batch 3: UI views and durability follow-ups.
    • UI view repairs plus five residual closures: U05-01, U05-m6, U09-m3, U09-m1 and U06-06.
    • Durability follow-ups:
      • A leftover seed journal quarantines only its own thread.
      • HeldRuntimeStore honours seed journals.
      • The lease sleep moves off async handlers.
      • Three probe-then-write writers become atomic.
  • Batch 4: Bun extension host lane, after its review blockers were fixed:
    • Node stays the default when the runtime is unset (D2).
    • Workspace and node_modules runtimes are never selected.
    • An override is final.
    • The runtime is pinned after the handshake.
    • A version mismatch is refused.
    • The Linux cap is clamped.
    • Batch 4 also pins the Pet artifacts eol=lf (the Windows CRLF byte-compare) and the Windows cfg for the conformance seam.
  • Batch 5: hosted-Linux repairs from wave b42 (PR CI 36728209612: 17792 passed; 17 failed):
    • A send with no model client is NotStarted again, so /edit restores the exchange it cut. That was a data-loss regression.
    • Host-managed admission waits for capacity, so a queued interrupt settles Interrupted instead of hanging InProgress.
    • The preview harness keeps its engine handle.
    • recorded_platform also replays os/shell.
    • Shell tool guidance renders the recorded shell during replay, not whichever shell first filled the process-wide cache (test: shared-process workspace gate fails on main while nextest CI passes #6698).
  • Batch 6:
    • TS Phase 0 remainder: generated host-protocol TypeScript, a lint that keeps core-only authority out of the host protocol, per-method cancel deadlines, and a visible host-down reason in /plugin and tool errors.
    • main through fix(tools): shell job retention, output deltas, and child process lifetimes #6759.
    • MCP auth classifier fix: an HTTP status is matched as a whole digit run outside URL tokens. The old bare "401" substring match read a connection reset on http://127.0.0.1:50401/mcp as an OAuth login requirement. That dropped the connection and flagged the server ◆ auth required.

Local evidence (macOS arm64, governed build slots)

  • Batch 5, broad selection (core::engine, runtime_threads, runtime_api, conformance, session_manager, tools::subagent, goal): 2350 passed; 1 failed.
    • The one failure is a local-only flake that hosted shared-process passes. See "Known" below.
    • Three production fix-off controls: 0 passed; 3 failed. Only the mutated case drifted.
    • Byte-exact restore, then 3 passed; 0 failed. Workspace Clippy passes.
    • Shell-guidance control: with the fix off, both golden suites drift on all 9 cases in 2 of 2 runs. Restored: 186 passed; 0 failed.
  • Batch 6 at this head:
    • Focused TUI (extension_host, approval, plugins, conformance, dependencies, mcp::oauth, mcp streamable HTTP, js_execution, process_tree): 190 passed; 0 failed.
    • Feature registry: 6 passed; 0 failed.
    • Node host: 63 passed, 0 failed, 1 skipped. Bun 1.4 host: 64 passed, 0 failed.
    • Typecheck passes, and the bundle rebuild leaves dist/ unchanged.
    • Four production fix-off controls (the params-object check, the method deadline, the host-down gate and the classifier): 0 passed; 4 failed. Byte-exact restore, then 4 passed; 0 failed.
    • Workspace CI-policy Clippy passes.
  • Pre-push checks: cargo fmt --check, git diff --check and the changelog mirror are clean; blocking-call budget 708 sites (within); dead-code 259/260; runtime-contract 55 metrics exactly at budget.
  • Public/private boundary: no security/* branch or private advisory commit is an ancestor of this head.

Hosted runs on this head

Known

  • Local-only flake. On this Mac, two shared-process runtime_threads tests time out on their 10 s watchdog. The turn monitor waits behind the process-wide test env barrier to resolve receipt secrets. Hosted shared-process passed both on b42.
    • This code path has been on main since 09-26.
    • A FIFO barrier made it worse and was dropped before any push.
  • Word signals such as "unauthorized" are still matched anywhere in MCP error text.

Still required before merge

  • Positive exact-head Linux, macOS and Windows tests and doctests on this PR.
  • shared-process-twice passing on the final SHA.
  • In-flight lanes:
    • D4 script tools
    • the Linux host sandbox plus startup/RSS gates
    • run_turn phase extraction
  • The remaining 0.10.1 inventory.
  • Final Linux VM and native macOS install, and real DeepSeek acceptance.

No release, tag, deployment or publication is performed by this PR.


Earlier body (b42e145), retained as dated history:

Current head: batch 2 (b42e145e46bedd6a40527823984eca84ba12c72b, tree d02f9048192e)

Fast-forward from the previous head 499dc6406ab2. It contains:

  • Phase 0 MCP consolidation. The legacy mcp-server aggregation proxy and the duplicate stdio client, pool and config writer are deleted (about 5,000 lines). The mcp-server spelling now aliases the existing in-process serve --mcp native server.
    • Hardening: strict JSON-RPC identity and params validation that never echoes request data; an initialize → notifications/initialized handshake that runs no tool before it completes, and none for id-less notifications; object-only arguments; a capped 16 MiB stdio frame reader with parse-error recovery.
    • Saved mcp.server_definitions data is retained, neither launched nor rewritten.
  • Hosted Linux repairs from wave 499.
    • Conformance goldens replay the platform posture they were recorded under (recorded_platform in each case). The runner's sandbox probe no longer leaks into <turn_meta>. No golden was re-recorded and no mask was widened.
    • The Codex OAuth pricing fixture now supplies the real ChatGPT endpoint.
    • The skills summary fits the 120-character registry limit.
    • The changelog mirror credit is synced.
  • main 568abae0: the exact-head green merges of fix(app-server): keep daemon threads, config, and bridge consistent across restarts #6772, chore(deps-dev): bump the npm_and_yarn group across 1 directory with 3 updates #6791, feat(providers): add Cheaper Inference as a bundled descriptor row #6761 and feat(web): move the community page onto the dictionary spine (#5337) #6794. Only documentation conflicted, and both sides were kept.
  • Terminal backpressure. It uses the existing 256-slot event queue reservation authority only. Start and terminal capacity are reserved before session mutation, and completed parallel tool results are kept (with their guarded projection) when the turn is cancelled.
  • Durability.
    • External rename, archive and delete hold the existing session lease for the whole mutation.
    • Seed history is journaled and committed last.
    • Goal status transitions compare-and-swap.
    • REPL rounds are bounded and poison on error.

Local evidence at this exact head (macOS arm64, governed build slots)

  • Focused Rust:
    • TUI/MCP: 251 passed; 0 failed; 1 ignored
    • runtime goal_loop: 20 passed; 0 failed
    • config descriptors: 10 passed; 0 failed
    • real CLI mcp-server alias binary: 2 passed; 0 failed
    • feature registry: 6 passed; 0 failed
  • Platform replay control. Pinning a Linux-shaped platform in two cases reproduced the exact hosted drift ("policy only; no execution sandbox available" and "sudo/setuid allowed (no-new-privs relaxed at startup)") in exactly those two cases. The fixtures were restored byte-exact, and conformance then passed 13 passed; 0 failed.
  • Phase 0 guard-off controls at 8d3b3f0, production file edited and tests unchanged:
    • Removing the handshake phase gate → 12 passed; 1 failed (a pre-handshake write executed).
    • Removing the notification-call guard → 12 passed; 1 failed at the canary-file assertion.
    • After byte-exact restore: 13 passed; 0 failed.
  • Terminal lane at 1194d0ad:
    • 37 passed; 0 failed.
    • Collector reverted to its pre-fix bytes → 36 passed; 1 failed at the completed-success assertion.
    • After restore: 37 passed; 0 failed, and Clippy passed.
  • Durability lane at 9413a891:
    • 154 passed; 0 failed; 1 ignored.
    • Released-lease control: 1 passed; 1 failed.
    • After restore: 2 passed; 0 failed.
  • Node/web: 799 passed; 0 failed (75 wrapper, 19 SDK, 54 extension host, 651 web). npm run check:web passed.
  • Workspace Clippy (the CI Lint policy): see the commit receipt.

Still required before merge

  • Positive exact-head Linux, macOS and Windows tests and doctests on this PR.
  • The shared-process-twice workflow dispatch for test: shared-process workspace gate fails on main while nextest CI passes #6698 on this exact SHA.
  • Remaining packets:
    • Bun host: blocked on its review findings (the default runtime must stay Node).
    • Durability follow-ups.
    • UI view residuals.
  • Final Linux VM and native macOS install, and real DeepSeek acceptance.

No release, tag, deployment or publication is performed by this PR.


Earlier body, retained as dated history:

The v0.10.1 candidate repairs live permission changes, idle task-store polling, scrolling and transcript noise, conversation undo/retry, nested-work deadlines and retained receipts, hook input, search recovery, and stream/transport configuration. It also incorporates reviewed MCP/Fleet/composer fixes, current Chinese copy and documentation, development dependency updates, and corrected cross-platform test fixtures. Existing PR histories remain intact and continue through their own exact-head checks.

Refs #6094, #6787, #6728, #6652, #6788, #6511, #6582, #6746, #6747, #6700, #6573, #6546, #6379.

Local evidence:

  • Combined all-feature CLI/TUI/config/protocol/localization and terminal-acceptance targets compiled. CLI103/0, configuration10/0, protocol6/0, localization52/0. The real automation-editor terminal scenario passed at five sizes, including save, restart and confirmed deletion.
  • The initial combined TUI selection passed239, failed one wrapped Chinese-copy assertion, and ignored two explicit measurement tests. A focused follow-up corrects only the comparison of text across layout rails; all41 follow-up checks passed, including real Node extension-host behavior, deferred first-tool calls, cleared to-do and stopship receipts. Across the two batches452 focused tests passed; two measurement tests were intentionally ignored.
  • Actual undo/retry flows cover Engine requests, retained compaction checkpoints, saving/reopening and eight ordered state races. Independent removal of the state guard and checkpoint preservation each failed at the intended runtime assertion; fixed-source14/0 passed again after byte-exact restoration.
  • Stream configuration proof includes41/0 plus7/0 restored checks and a real loopback HTTP/2 PING/ACK-timeout observation. Disabling canonical precedence, saved-value precedence or actual keepalive application each caused its regression to fail.
  • Seven earlier deadline/hook-receipt negative controls failed and restored positives passed. Hook input includes actual failed shell execution with foreground/background observers; nested evidence is saved with the enclosing tool result, not claimed durable before handback.
  • Fresh combined npm774/0 (75 wrapper,16 SDK,54 extension host,629 web), web checks passed. VS Code42/0 and compile passed. Final notes18/0, contributor credit8/8 across three surfaces,35 feature references and packaged changelog sync passed.

These are source, local fixture and bounded terminal results. Final Linux/Windows/macOS CI, two complete shared-process workspace runs at one SHA, optimized-candidate CPU comparison, native acceptance, and release authorization remain separate. The optimized build is not an installation or release. Candidate notes retain explicit Linux no-new-privileges and nested-work durability limits. Bun/TypeScript host expansion and broader post-GA features remain outside this candidate. No release, deployment, tag, credential change or publication is performed by this PR.

🤖 Generated with Claude Code

Hmbown pushed a commit that referenced this pull request Sep 30, 2026
…indows-safe

- crates/tui/CHANGELOG.md regenerated with scripts/sync-changelog.sh
  (Version drift on #6782 failed: "crates/tui/CHANGELOG.md is out of date
  with the root CHANGELOG.md slice"); `sync-changelog.sh --check` now passes.
- skills::package_digest::tests::bounded_read_accepts_exact_remaining_bytes
  used the anonymous tempfile::tempfile(), which Windows CI denies under the
  hermetic test home ("Access is denied", os error 5, seen on #6783's
  Windows job 109657004556). It now uses a named file in a tempdir, like the
  neighbouring test. rustfmt --check clean; Windows CI is the proof.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Hmbown and others added 29 commits September 29, 2026 20:04
Six typed-save fixtures used quiet even though the validated enum accepts normal, concise and verbose. Use concise for these writes; preserve the separate raw legacy-load coverage. Six focused config tests passed, zero failed. Shared source gate: npm test674 passed/0 failed; check:web passed. No production configuration behavior changed.

Signed-off-by: Hunter B <hmbown@gmail.com>
… clearance

Third review of GH6787 found regressions versus v0.10.0 for interpreter-embedded deletes, earlier path rebinding and wrappers that reinterpret absolute paths. Only a direct literal rm may use pre-execution workspace clearance. Preserve ordinary relative/absolute workspace cleanup. Move policy tests into their UI owner instead of raising the core-to-UI boundary budget; keep actual child execution coverage. Fix a Unix-only test fixture declaration for Windows.

Evidence:26-case v0.10.0/wave/fixed source comparison; two isolated tests pass and fixes-off control fails as expected. Repository command_safety tests78 passed/0 failed. Combined focused TUI tests143 passed/0 failed, including workspace cleanup, detached holds and Full Access child execution. Config fixtures6/0; tool guidance2/0. npm test674/0; check:web passed. Boundary scan passes with baseline lowered, rustfmt/diff checks pass. Linux native acceptance and final hosted CI remain open; no release claim.
Signed-off-by: Hunter B <hmbown@gmail.com>
Preserves the previous stopped session measurement harness before takeover. No implementation was present in this lane. Not compiled or run; not a release-ready change.

Signed-off-by: Hunter B <hmbown@gmail.com>
…transcript (#6652)

UNVERIFIED: not yet compiled or tested after the change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Signed-off-by: Hunter B <hmbown@gmail.com>
Refs #6652. Pending wheel input and scrollbar drag now select interactive
cadence. Preserve transcript cache on actual height-only resizes, including
long final answers. Cache collapsed-tool summaries and direct lookups once
per history/active/expansion generation. Preserve queued input ordering and
include resize/refit clears inside the DEC 2026 synchronized draw.

Verification on macOS arm64, focused all-features TUI tests:
- Governor build + g3_ filter: 7 passed; 0 failed; 1 ignored (six new G3
  regressions plus one existing name match). Linker warns about a large
  unwind table; the test executable links and all selected tests pass.
- Transcript/cache/group/todo focused follow-up: 50 passed; 0 failed;
  2 ignored. Reused the same executable, no second compile.
- Four ignored benchmark/measurement harnesses: 4 passed; 0 failed.
- rustfmt on owned files and git diff --check passed. npm/web gates were
  not run for this Rust-only slice; hosted CI and Windows remain unverified.

Unoptimized TestBackend, 200 frames, width 140:
- 400 cells: 3.817 ms scroll / 3.545 ms height change per frame.
- 4000 cells: 5.129 ms scroll / 6.402 ms height change per frame.
- Isolated 4000-cell/1000-group lookup: prior linear search 14.796 ms,
  cached direct lookup 0.464 ms per pass. At 400 cells/100 groups: 96.94 us
  versus 103.91 us. This compares lookup algorithms in the same binary,
  not whole-frame before/after or actual terminal throughput.
- Retargeting 17999 lines: 400 retargets in 3.789 ms, 798 rows reflattened.

Windows input-to-paint latency, terminal byte throughput and release-build
performance remain unverified. Frame bookkeeping still scans history;
width changes still reflow it. No #6652 auto-close claim.

Signed-off-by: Hunter B <hmbown@gmail.com>
Keep streaming thought to a three-row tail, settled calm thought to one
localized duration header, and failed Generic/MCP output to a six-row
head/tail excerpt. Full details and explicit reasoning expansion survive.
Place the reasoning reveal affordance in the header; preserve copyable
preview content and existing focus ownership. Mute successful tool chrome,
remove the repeated MCP name, and expose group count/reveal affordances.
Reuse one six-row cap for calm and hidden tool details.

Source checkpoint for the coordinated shared build. Rustfmt and diff checks
passed. All 15 complete locale packs parse with both new typed messages and
matching duration placeholders. Prepared 4x3 frame and tool-state matrices,
plus focus, copy, full-output, style and cap regression coverage.
Rust compilation, focused tests and rows-per-turn measurement are pending
the shared build slot; no npm/web gate or hosted CI is claimed here.

Refs #6652

Signed-off-by: Hunter B <hmbown@gmail.com>
Adapt copy and offscreen-target fixtures to the three-row live tail and settled header. Keep the explicit expand through a streaming-to-settled transition. Rustfmt and diff checks pass; runtime qualification remains queued with the combined TUI build.

Signed-off-by: Hunter B <hmbown@gmail.com>
Refs #6728. A settled readable queue fingerprint now removes the empty-claim
fallback deadline. Admissions and queue changes still wake workers; unreadable
or coarse recent metadata and failed claims retain bounded retries. Claim locks,
execution leases, cancellation observers and startup crash recovery are unchanged.

Strengthen idle reload assertions to exactly zero after settling. Add regressions
for retrying unknown/recent metadata, external queue writes without a local
notification, and running-task cancellation with unchanged queue metadata.

Checkpoint evidence: rustfmt --edition 2024 --check and git diff --check pass.
Exact extracted ClaimSchedule with the new regression tests: baseline 1 passed,
1 failed (unchanged queue incorrectly becomes due); fixed 2 passed, 0 failed.
The clean ac9ce1f release baseline build passed in 27m04s. Fresh-HOME,
460-task, 60-second before/after measurements and patched release build follow.
The real task_manager:: tests will run in the parent's combined TUI test build.
The npm test and check:web gate is owned by the parent integration lane; no pass
for those or hosted CI is claimed by this checkpoint.
Unqualified checkpoint of ten inherited dirty paths before review. The conversation actions contain an apparent fixes-off negative-control state; restore and validate the complete behavior before integration. No passing tests or release readiness claimed.

Signed-off-by: Hunter B <hmbown@gmail.com>
The recovered GH6788 patch had its fixes disabled and used a generic action
sequence that could continue after a rejected sync. Conversation undo now
prepares the desired history without erasing the live transcript, selects
the last real user boundary through the existing runtime classifier, and
uses one dedicated UI action for undo and retry. That action refuses busy
or changed conversations, obtains Engine acknowledgement, then requires a
CompletedCommit/FlushAndReport durability receipt before replacement
inference. Failed save/open/channel operations report an error and never
send the retry. A save failure leaves the acknowledged undo visible.

The inherited Engine-to-UI regression was replaced with a UI-owned test
that calls the real apply_command_result path, captures mock provider
requests, reopens the saved session, checks checkpoint retirement, and
exercises runtime ownership, filesystem save failure and closed Engine.
Original file undo behavior remains on the existing SyncSession route.

Validation: rustfmt and git diff --check pass; command migration manifest
PASS. Blocking-call ratchet reports existing untouched auto_review.rs
canonicalize 3>0 and feature_registry.rs std_fs 1>0; no new slice sites.
No Rust build/tests or npm test/check:web run in this checkpoint: the
coordinator owns shared build slots and web gate. Integration and focused
regressions are prepared, not claimed passing. No push/merge performed.

Refs #6788
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Hunter B <hmbown@gmail.com>
Extend the UI-owned GH6788 fixture to restore the actual saved record through the loaded-session projection and SyncSession apply action, then send another turn through apply_command_result. The captured mock model request must contain only the retained prompt and the new prompt, with no undone prompt or answer. File LoadSession itself always replaces the injected client, so the fixture explicitly exercises its existing projection with the injectable Engine.

Validation: rustfmt and git diff --check pass. This Rust fixture remains uncompiled and unrun pending the coordinator build slot; no provider, native, npm or hosted CI qualification claimed. Existing save failure and closed Engine assertions still forbid additional inference.

Refs #6788

Signed-off-by: Hunter B <hmbown@gmail.com>
Existing guard/history/exhaustion repairs remain intact. Stamp nested execution from the Engine clock after approval, carry its absolute deadline through llm_query and every recursive RLM, and bound inline Python plus forwarder handback by it. Expired work is refused before dispatch. Timed-out or broken persistent Python state is discarded, and the event forwarder aborts with its owner.

Collect nested code/status events once at their producing loop in the existing shared RLM receipt batch. Include the receipts in normal and kernel-error tool bodies so all hosts retain them through existing tool-result/session/artifact persistence. Existing64 provider-call slots plus64 nested-run slots bound work and repeated setup failures; there is no new event loop or store. In-flight receipts can still be lost if the whole tool future is externally dropped or the host is killed before result handback; this is stated beside the owner, not claimed crash durability.

Prepared tests cover spent/expired Engine budget propagation, pending plain/recursive calls, receipt admission bounds, exact-once live forwarding, Python refusal before a filesystem marker, and loopback-provider nested execution whose success/error tool records survive SessionManager save/reopen. Negative controls remain pending: reset recursion to the default deadline; remove event collection or the result nested_events field; remove expired-parent preflight. These must respectively fail timeout, receipt replay, and no-side-effect assertions.

Validation: rustfmt, git diff --check, command-crate boundaries and command migration manifest pass. No Cargo test/build, native/provider, npm gate, hosted CI, push or merge qualification for this slice. Coordinator owns the next combined build.

Refs #6511

Signed-off-by: Hunter B <hmbown@gmail.com>
Retry dispatch deliberately writes a new in-flight checkpoint. Flush and inspect its retained prompt plus one retry instead of expecting no checkpoint after replacement inference. Seed and flush the old checkpoint first, and require standalone undo to clear it. Remove the obsolete pop_api_message production helper and use the existing truncate helper in its sole fixture consumer.

Validation: rustfmt and git diff --check pass. Root combined run exposed the old checkpoint assertion failure and default-feature dead_code error; repaired Rust tests and default build are pending the root combined batch. Refs #6788.

Signed-off-by: Hunter B <hmbown@gmail.com>
Preserve distinct tool glyph shapes while expecting the admitted muted success ink and group reveal cue. The old 60x8 focus fixtures now fit both compact cells; keep newest-visible ownership there, then assert a real one-row 60x5 viewport before checking pending-scroll and older-reasoning ownership.

Evidence: root combined batch compiled and ran 437 passed, 6 failed, 3 ignored; five failures identified as stale Calm expectations, with undo fixture owned separately. rustfmt and git diff --check pass. This corrected test checkpoint awaits one combined recompile; no passing Rust result is claimed for it yet.
Signed-off-by: Hunter B <hmbown@gmail.com>
Bound both registry and session lock acquisition by the inherited absolute deadline. Inline Python startup, context refresh, and every code round share one parent-clamped deadline regardless of whether a bridge is present. Preserve the existing error and broken-kernel retirement paths.

Validation: rustfmt and git diff --check pass. Added parent_deadline_bounds_rlm_session_lock_waits for both contended locks, with an outer timeout and no-side-effect assertion. Cargo execution is pending root combined build; no runtime pass claimed. Refs #6511.

Signed-off-by: Hunter B <hmbown@gmail.com>
Use the existing PythonRuntime String error in the compatibility eval timeout, matching inline execution. Static signature review, rustfmt and diff check passed; combined Cargo pending.

Signed-off-by: Hunter B <hmbown@gmail.com>
Checkpoint the 68 inherited translation paths before reconciling against current main and Calm locale additions. No source prose changed during recovery. JSON/docs/web/Rust validation and editorial review remain pending; this checkpoint is not qualified for landing.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Hunter B <hmbown@gmail.com>
Correct revoked-access meaning, next-turn behavior, provider routing and
Traditional Chinese wording; preserve the inherited contributor translations
and add the candidate's Calm thought keys. Sync hook stdin receipts, trusted
skill loading, MCP startup timeout, retry/HTTP settings, question sheet and
canonical runtime job paths. Explain deliberate Auto-Review questions and
indefinite default question waits; remove unsupported Windows confinement and
fresh hardware-validation claims. Replace the stale glossary census with a
usable Chinese guide.

Local evidence: web i18n 56 passed/0 failed (7 files); TypeScript exit 0;
check:docs and check:locales passed; local GT export unchanged. Candidate
English/TUI parity: both Chinese packs 2434/2434, exact placeholders and changed
edge whitespace. Parsed 315 local links across 46 docs: 0 failures; 4327
technical identifiers across 43 source translations present. diff --check
passed. G1 read-only safety review found no additional Chinese meaning drift.

Full npm test && npm run check:web and Rust localization tests remain the
parent's combined-candidate gate; not run or claimed in this lane. Native
Traditional Chinese, layout, Windows execution, provider, hosted CI and release
qualification remain separate. No push or release.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Hunter B <hmbown@gmail.com>
…tracts

Portable imports reject machine-bound authority settings. Keep closed-choice
acceptance coverage on portable verbosity and preserve unknown existing fields.
Run ordinary typed plugin activation under normal supervision rather than the
600 ms fault watchdog, and use the clippy-approved search error assertion.

Validation on combined source 69824d3 plus recorded patch: CLI bundle tests
61 passed, 0 failed; typed plugin activation 1 passed; custom search fallback
1 passed. Combined all-feature CLI/TUI/workflow/workflow-js test build passed.
Unrelated Calm glyph assertion remains open (TUI combined 97 passed, 1 failed).
Full npm test: 740 passed, 0 failed (68 wrapper + 16 SDK + 54 extension + 602 web);
check:web passed. Windows handshake scheduling remains separate, unverified.

Signed-off-by: Hunter B <hmbown@gmail.com>
Project the existing bounded local-shell execution receipt into the requested
tool_call_after JSON envelope for both direct and queued observers. Preserve
legacy environment fields and refuse unsupported, incomplete, or unknown
execution identity. Report actual command/cwd, nullable exit status, separate
output streams and truthful truncation flags; never reconstruct from intent.
Observer output cannot deny or rewrite the completed tool.

Validation: 3 new hook tests and 5 existing receipt/shell tests passed in the
combined all-feature TUI binary. The real shell fixture exits 7 and reaches
both synchronous/background and direct/queued observer paths. Removing stdin
dispatch makes that regression fail; restored source passes. Six RLM deadline
and receipt negative controls also failed as intended and passed restored.
All four mutated production files restored byte-identically. Full npm gate
740 passed, 0 failed; check:web passed. Native external MemoryWhale end-to-end
qualification and non-macOS execution are not claimed. Fixes #6582.

Signed-off-by: Hunter B <hmbown@gmail.com>
The Calm rail inserts span 0 and the shared status mark is span 1; tool-family identity is span 2. Correct the regression to inspect the actual family glyph without changing its identity or quiet color checks. Exact test passed 1/0 in the rebuilt binary; other five repaired Calm/undo regressions passed in the preceding combined run. npm740/0 and check:web passed.

Signed-off-by: Hunter B <hmbown@gmail.com>
…h runtime

Auto-Review still exposes deliberate questions in interactive hosts; omitted or zero question timeout is indefinite. Remove a stale provider-count claim and update two Chinese wording expectations to the reviewed current-session labels. Combined localization tests52/0; full npm740/0 and check:web passed. Command boundary and migration checks passed. Chinese native layout and real provider qualification remain separate.

Signed-off-by: Hunter B <hmbown@gmail.com>
CodeWhale Bot and others added 29 commits September 30, 2026 07:25
… 98b1339b, fd54c004)

These are the ui-views lane's three WIP checkpoints, squashed with their authored
dispositions. The old lane merge e977f624 was not carried; this applies only the
lane's own changes onto the current integration.

- U06-06: aplay/ffplay resolve only from fixed absolute prefixes, never PATH,
  with the exec bit checked; the macOS companion launcher is /usr/bin/open.
- U05-01: a settled workflow run is never reopened by a late TaskStarted, and a
  late new row in a cancelled run is finalized Cancelled.
- U05-02: a run_completed/task_completed without `status`, or still saying
  running/pending, fails closed instead of reading as success.
- U05-04: delegate and fanout cards never let later envelopes change a terminal
  card or slot.
- U05-05: a follow-up receipt is appended only to the focus on its named agent
  (or its fork target).
- U05-m1/m2/m4/m6: phase routing destination hardening; a settled fanout
  reports its worst outcome; only bare Left/Right switch dock tabs; the fanout
  grid wraps within the render width.
- U06-03: pet live batches are bounded by the owner's request byte limit as
  well as the 64-event count.
- U06-m4 (partial): the export wait is aligned with the owner's work bound.
- U09-03/05/m2/m3/m5/m6: fleet model drafts install only unchanged answers;
  extension tabs re-anchor selection by entity key; settings Tab/BackTab
  preview; automations and agent projections keep recorded errors and reasons;
  fleet list and wrapped detail views scroll by painted rows.
- U09-01 (Refs #6555): fleet detail Save and rename refuse a source changed on
  disk since it was opened.
- U08-09 (Refs #6555): the session picker lists the store off the UI thread and
  shows list errors, never an empty store.
- A late Started for a recovery card names it without reopening it.

Lane receipts, dated: an earlier broad run gave 763 passed / 1 failed (a
completion-before-started recovery case). The fd54 correction then passed a
52/0 subset. No whole-set exact-head pass is claimed; this integration's
qualification is the proof.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
U05-01: a repeated or replayed `run_started` for the run a panel already
shows no longer rebuilds it, which erased settled rows and the terminal
outcome. It names the run and fills missing metadata (start time, budget,
source). The guard sits in `apply_event`, so the live `app.rs` path is covered
as well as the JSON path. The runtime emits `RunStarted` once per fresh run
id, so a same-run start is only ever a duplicate.

U05-m6: the fanout header fits 0-3 columns. It shows the glyph only when it
fits and the count only when there is room after it, and a zero-width render no
longer chunks the dot grid.

U09-m3: a spawned-but-not-started worker is projected with
`worker_status: Queued`. The /subagents row reads "queued" and the whale state
shows Thinking instead of Running. No new `SubAgentStatus` variant (663
serialized uses).

U09-m1: the settings view keeps the committed editor until the host answers.
A rejected value reopens the editor with what the user typed, on the view
rebuilt from disk truth. Accepted values still close it.

U06-06: the Linux pet browser launcher is `xdg-open` from a trusted system
prefix, never `$BROWSER` or PATH, so a workspace cannot shadow it and receive
the owner token in the URL. It is reaped off the pet worker thread. The
trusted-prefix resolver is renamed `trusted_system_executable` for its three
callers. macOS and Windows keep OS default-browser resolution.

Remaining, recorded beside the code: U06-m4 (export blocks the QuickJS world
until incremental export exists in pet-native.js) and U09-01/#6555 (fleet save
is check-then-rename, not a locked compare-and-swap).

Tests added after the fix: a_repeated_start_of_the_same_run_keeps_settled_rows_and_outcome,
fanout_header_fits_zero_and_one_column_renders,
a_rejected_setting_reopens_the_editor_with_the_typed_value,
a_pending_worker_row_reads_queued_not_running. Compilation and fix-off controls
run in the integration batch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
The test module reached json! only through a glob of its parent's private
import; the explicit path does not depend on that.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
The closure's early returns go through `?`; naming its Ok error type keeps
inference from depending on the handler's tail expression.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
… Node once

- S1: the manager pinned the resolved runtime before the host was spawned,
  so a Bun that resolved but could not start stayed pinned for the session.
  The runtime is now pinned only once a host on it completes the handshake
  (`RuntimePin`, one mutex, never held with another lock). Under
  `runtime = "auto"`, before anything is pinned, a Bun host that fails to
  launch or handshake is reported in one diagnostic naming the Bun, its
  path and why, and Node is resolved for the rest of the session (the
  summary keeps `runtime = "auto"` and says Bun failed to start). The Node
  attempt runs under a fresh host generation, so the failed Bun host's late
  exit callback cannot be taken for the Node host's; a superseded start
  still aborts. An explicit `bun` or `node` never falls back. Launch
  preparation moved into `prepare_launch` (blocking, under spawn_blocking
  as before). With no successful handshake nothing is pinned, so `/plugin`
  shows no runtime line after a failed first launch; its failure reason is
  in the status and diagnostics.
- S4: deleted the post-handshake "did not apply its kernel memory limit
  (its stderr says why)" check. After a successful handshake the host's
  enforcement is either the requested jetsam limit or `launch.memory`,
  which `plan_launch` sets to `MemoryEnforcement::planned(kind)`, so the
  check could not fire in production and its message guessed a cause.
  `HostLaunch.memory` stays: it is the seam the handshake test uses to
  request a jetsam limit a Node host cannot apply. `render_status` no
  longer recomputes planned enforcement from the pinned runtime; it
  describes the running host's settled enforcement (`HostProcess::memory`
  via `HostStatus::Ready`) and nothing when no host runs.
- F6: the macOS operator-kill assertion checked for "exceeds its", which
  nothing emits; it now checks the real heartbeat message "exceeded its
  memory cap".
- B2: the module's known-limitations note now says Node is the default,
  Bun is an opt-in qualified on macOS only, no Bun host has run on Linux or
  Windows, and the native-code lockdown covers the entry points found so
  far.

New unix test `auto_uses_node_for_the_session_when_the_bun_host_fails_to_start`:
a fake `bun` that passes the 1.4.0 version probe but exits before
`host/hello`; under `auto` the host ends Ready on Node with exactly one
fallback diagnostic, a Node summary that keeps `runtime = "auto"`, and two
spawn attempts; under `runtime = "bun"` it ends Failed after one attempt
with nothing pinned.
Not compiled here (Rust build slots are governed); rustfmt --check clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…ed external writes)

Independent-review findings, verified against source before fixing:
- A leftover or uncertain seed journal quarantines only its thread (records and
  journal untouched, held out of recovery) instead of refusing the whole store.
  A journal left behind after a commit, with later turns above a complete,
  owned seed, is judged committed.
- HeldRuntimeStore reconcile judges seed journals the same way (read-only), so
  an unpublished seed can no longer reach recovered history through reconcile.
- Runtime API session patch/delete and lease reservation run under
  spawn_blocking. Export, save and scrub-secrets hold the session's live lease
  across load and save (409 when open interactively; scrub reports it busy).
- A malformed id on delete is 400; an unknown id is 404 without creating lease
  files.
- A REPL kernel that a dropped turn left broken is replaced by the next turn.
- Known limitations are written beside the code: claim_live_session is
  best-effort, SSH confirmed stop is local only, the plugin catalog path check
  is textual, and stale-checkpoint pruning still uses a released probe.

Qualification is in the batch 3 receipt.

Signed-off-by: CodeWhale Bot <bot@codewhale.net>
F1: `run_doctor` is async and resolved the extension-host runtime inline,
which runs every candidate `bun`/`node` (`--version` and, for Node, one
flag probe per native-code flag) on a Tokio worker. It now resolves under
`tokio::task::spawn_blocking`, as the host launcher already does, and a
probe task that does not finish is reported instead of panicking.

Known limit, written beside the call: the probes have no timeout, so a
runtime binary that hangs on `--version` stalls doctor there. The codebase
has no bounded-output command helper to reuse; the `wait-timeout` crate is
a dependency (used by shell/hooks) if a timeout is wanted later.

Blocking-call budget: `crates/tui/src/dependencies.rs` gains `std_fs: 2`
for the two `canonicalize` calls in `untrusted_location` (6d114f3). They
run only inside `resolve_extension_host_runtime`, which is documented
blocking and whose production callers (the host launcher and now doctor)
both call it under spawn_blocking; the ratchet cannot see through the call.
Only that entry was added by hand: `--update` would also have rewritten
unrelated entries (it drops lib.rs and last_round.rs and lowers voice.rs),
which is not this slice's to change.
`python3 scripts/check-blocking-calls-budget.py`: "713 sites across 210
files, within budget". Rust not compiled here (build slots are governed);
rustfmt --check clean on lib.rs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…odule

The views test module imports names explicitly, so the queued-worker test
names super::lifecycle_worker_status and super::format_agent_status.
Found by the batch 3 compile (E0425 x4).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
B2: the changelog, extension docs, design doc, example config, CI comment
and code comments said Bun is the default runtime with Node as the
fallback. The governing decision keeps Node the default and diagnosed
fallback until the four Bun gates are proven on every platform and a
cutover is recorded; they are proven on macOS only. Every one of those
statements now says: Node is the default; Bun is an opt-in through
`runtime = "bun"` or `"auto"`; `auto` prefers a supported Bun and uses Node
when none is found or the Bun host fails to start (S1).

Also corrected while there:
- CHANGELOG (0.10.1 unreleased): states the known limit that the
  native-code lockdown covers only the entry points found so far, and no
  longer implies kernel enforcement on Linux/Windows was proven for Bun:
  the Rust host tests, the memory cap included, have not run a Bun host on
  Linux or Windows (CI runs them with Node; the JS suites run under Bun
  1.4.0 on Linux). It also records the override, PATH-skip, pinning and
  version-refusal behaviour from this series.
- docs/EXTENSIONS.md claimed neither runtime reads `tsconfig.json`. Checked
  2026-09-30 with Bun 1.4.0 and Node 26.10 through the real host bundle
  (test harness `startHost` + `activate`): a plugin importing
  `@alias/thing.mjs` activated under Bun (`{"status":"ok"}`) via a
  `tsconfig.json` `paths` entry next to the entry file, and also via one in
  the directory above it; under Node it failed with "Cannot find package
  '@alias/thing.mjs'". The doc now says Bun applies a nearby tsconfig's
  `paths`/`baseUrl`/JSX settings, Node does not, and records as a known limit
  that a tsconfig above the reviewed bundle can steer a Bun host's imports.
  The Limits section no longer says the cap is kernel-enforced for every
  host (a macOS Node host, the default there, is heartbeat-checked).
- docs/design/TS_EXTENSION_HOST.md: the "As built: Bun runtime" section,
  its selection/pinning/test notes, the SIGKILL exit wording and the
  handshake sketch (`node_version` -> `runtime: {name, version}`).
- config.example.toml and ci.yml comments; runtime.ts header (comment-only:
  `npm run typecheck` clean, `npm run build` leaves dist/ byte-identical,
  `git diff --quiet -- dist` exit 0).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…-exec

S3: on macOS a Bun host re-executes itself in place to apply its jetsam
limit, rebuilding argv as `[process.execPath, ...process.execArgv,
...process.argv.slice(1)]`. If execArgv did not carry `--no-install
--no-env-file --config=/dev/null --no-addons`, the host that loads plugins
would run without them.

Observed 2026-09-30, macOS 26.1 arm64, Bun 1.4.0, with the real bundle
through the test harness (`startHost({ env: {
CODEWHALE_HOST_MEMORY_LIMIT_MIB: '300' } })`, then a plugin tool returning
`process.pid` and `process.execArgv`): hello reported
`memory_limit_mib: 300`, the pid after the re-exec equalled the spawned
pid, and execArgv was exactly ["--no-install","--no-env-file",
"--config=/dev/null","--no-addons"]. Plain `bun --no-install --no-env-file
--config=/dev/null --no-addons probe.mjs` also reports those four in
execArgv. So no code change: the existing macOS test's `pid` tool now
returns execArgv too and asserts it equals HOST_ARGS after the re-exec.

Negative control: with dist's re-exec line temporarily changed to drop
`...process.execArgv`, `bun test ./test/host.test.mjs -t "kernel memory
limit"` failed (0 pass, 1 fail, AssertionError at the execArgv deepEqual);
dist restored (`git diff --quiet -- dist` exit 0).
Suites on this tree: `npm run test:bun` (Bun 1.4.0): 63 pass, 0 fail;
`npm test` (Node 26.10.0): tests 63, pass 62, fail 0, skipped 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…plexity

Batch 3 workspace Clippy (-D warnings) rejected the pending listing cell's
nested type in the U08-09 session picker. Name the result once and use it at
the cell, the landing function and the store scan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…aller exists

Wave b42e145 hosted Windows failed `cargo test --no-run` with
`method pin_recorded_platform_posture is never used` under -D warnings: the
scripted events/prompt/hooks conformance families are `#[cfg(unix)]`, so on
Windows the cfg(test) seam had no caller. Gate it with the same condition.
Unix builds and the replay behaviour are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Pet conformance `recorder (windows-latest)` on wave b42e145 reported
pet-native.js (TUI and iOS copies) and ios demo.jsonl as stale. 1fceabf
made `npm --prefix pet run check` compare the regenerated LF bytes against
the committed files, which `* text=auto` checks out as CRLF under Windows
autocrlf. No commit in the range changed a pet input, and the last Windows
pass predates that check. `owner.rs` also embeds pet-native.js with
include_str!(), so a Windows build would embed different bytes. The repo's
existing rule for byte-compared and include_str! inputs is `text eol=lf`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…ays the default)

Lane 45e3d3a plus eight review-fix commits, reviewed against source by the lead:
- B1: an unset runtime means Node; a Bun-only table means Bun; auto and bun are opt-ins.
- B2: docs, CHANGELOG and comments stop claiming a Bun default or unproven
  Linux/Windows enforcement.
- S7: a searched runtime inside node_modules or the working directory is skipped
  with a recorded reason.
- S8: a configured override is final and fails loud.
- S1: the runtime is pinned only after a successful handshake; under auto a Bun
  that fails before any pin falls back to Node once.
- S2: a runtime version changed mid-session is refused.
- F2: the Linux RLIMIT_DATA cap is clamped to the inherited hard limit.
- F11: NODE_OPTIONS is blank for the host.
- F1: doctor probes run under spawn_blocking; the missing timeout is a recorded
  limit.
- S4: the dead mismatch diagnostic is deleted.

Lane JS: Node 62/0 (1 skipped), Bun 1.4.0 63/0. A real control shows the macOS
re-exec keeps the four lockdown flags. Known limits recorded: Bun reads
tsconfig paths above the staged plugin, and there is no Linux/Windows Bun
host run. The extension_host feature stays default-off and Rust owns turns.
Rust qualification is in the batch 4 receipt.

Signed-off-by: CodeWhale Bot <bot@codewhale.net>
scripts/sync-changelog.sh --check reported the mirror out of date after the Bun
lane merge corrected the root CHANGELOG's runtime-default wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…d host turns

Hosted Linux on wave b42e145 (run 36728209612) failed 15 tests. These two
integration regressions came from the terminal backpressure lane and reproduce
on macOS (local 44 passed / 15 failed; bisect: all four sampled failures fail at
lane head 1194d0a alone):

- A send that never reached a model client returned Finished(Failed), not
  NotStarted. `/edit` restores the exchange it cut only for NotStarted
  (C02-02), so an edit whose replacement could not dispatch silently dropped
  the user's original prompt and answer. Goal reconciliation also read a
  failure instead of "could not start". The reserved TurnComplete is still
  settled through the permit; only the outcome returns to NotStarted.
- Admission refused a queued, already-cancelled turn with no lifecycle. For a
  durable host (Runtime threads, host_managed_turns) that had already recorded
  the turn, nothing ever terminalized it: it stayed InProgress with an active
  claim. Host-managed admission now waits for capacity instead of refusing on
  cancellation, so the turn emits TurnStarted and settles Interrupted with zero
  provider calls, which is the runtime contract. Interactive queued
  cancellation keeps the lane's no-fabricated-lifecycle rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…tures

wire_preview_engine dropped the EngineHandle on return, so every preview turn
ran against a closed event channel. Since the backpressure lane, admission
refuses a turn whose event consumer is closed ("Cannot start the turn because
its event consumer is closed"): no production embedding runs a turn nobody can
record. Ten preview wire-body tests then saw zero provider calls. Return and
hold the handle, as every production embedding does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…onment block

With the sandbox posture replayed, hosted Linux reached the next host fact:
events `prefix_cache_change` carries "frozen: <hash>" of the system prompt
prefix, whose `## Environment` block names the OS and shell (macOS/zsh when
recorded, linux/bash on the runner). Each case's `recorded_platform` now also
records `os` and `shell`. The harness replays them through a test-only,
thread-scoped override held for the scripted turn; the engine runs on that
thread's current-thread runtime. Production always reports this host. No
golden re-recorded; no mask added or widened; the prompt family's documented
`- platform:`/`- shell:` masks are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Phase 0 (CURRENT_DECISIONS §26): the extension-host protocol is now single-
sourced in crates/tui/src/extension_host/protocol.rs.

- protocol::METHODS is the whole surface (name, direction, request or
  notification). Both parsers admit only table rows; `admit` replaces the
  per-arm expect_request/expect_notification helpers.
- The wire types derive schemars::JsonSchema under cfg(test). A new test,
  extension_host::protocol::tests::typescript_protocol_is_generated_from_the_rust_types,
  renders crates/tui/extension-host/src/protocol.generated.ts (constants,
  error codes, method table, every params shape the host validates, and the
  wire types) and compares it through the conformance golden helper
  (CODEWHALE_CONFORMANCE_UPDATE=1 re-records; refused under CI).
- protocol.ts drops its hand-written PARAMS table, method lists, interfaces
  and constants and validates against the generated shapes; root.ts and
  rpc.ts drop their restated ActivateParams/ActivateResult/deactivate result
  and MAX_INFLIGHT. Strictness now follows the Rust type's
  deny_unknown_fields exactly (as before for every host->core type).
- Rust params() refuses non-object params, so `host/ready` with `[]` no longer
  decodes positionally into EmptyParams (the TS side already refused it);
  corpus case 36 pins that for both sides. host/ping and host/shutdown decode
  as EmptyParams in the test-only core parser, matching the generated shape.

Evidence (this checkout; cargo not run here, the lead compiles):
- npm run typecheck: clean
- npm run build: dist/ rebuilt and committed
- npm test: tests 63, pass 62, fail 0, skipped 1 (Bun-only test)
- npm run test:bun: 63 pass, 0 fail
- node probe of dist/protocol.mjs: ready [] / null rejected, ping {x}
  rejected, activate extras accepted, hello memory_limit_mib null accepted,
  kind `command` and event/emit rejected

Known limit: the generated file was rendered by hand from the generator's
format; if the Rust test reports drift, re-record it, rebuild dist/ and rerun
the JS suites.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Phase 0 (CURRENT_DECISIONS §26): a test, run by the normal
`cargo test -p codewhale-tui --lib`, that the extension-host protocol never
gains event, store, approval, secret/credential/token/auth, turn-loop,
session or prompt methods.

extension_host::protocol::tests::host_protocol_never_gains_core_authority
checks three things over protocol::METHODS, the table both parsers admit
from (and the TypeScript validator's generated copy):
- no method name, in either direction, contains a core-only authority word;
- the table equals REVIEWED, where each current method carries the reason it
  gives the host no core authority (host->core requests registry/register and
  registry/unregister included), so any addition is a reviewed edit;
- every `ns/name` string literal in protocol.rs is a table row, so nothing is
  decoded or sent outside the table (40 literals today).

Fix-off controls for the lead: add `row(Direction::HostToCore,
"event/emit", false)` to METHODS (fails the word check); drop a REVIEWED row
(fails the set check); add a `"store/put"` literal to protocol.rs (fails the
source scan).

Evidence: rustfmt --check clean; cargo not run here (the lead compiles).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Phase 0 (CURRENT_DECISIONS §26), per-method cancel deadlines.

Already present before this change, under other names: `$/cancel` both
ways (protocol.rs cancel_value; the host aborts the handler in rpc.ts),
request_with_deadline cancelling ext/activate (ACTIVATE_DEADLINE + 1 s) and
ext/deactivate (DISPOSE_DEADLINE + 500 ms) at their call sites in mod.rs, and
tool.rs's own 120 s timeout + CancelOnDrop. The gaps: no single table, each
margin chosen at its call site, and an unbounded HostProcess::request() used
for host/initialize (bounded only by the outer handshake timeout) and test
shutdown.

Now:
- CoreRequest::deadline() is the table, an exhaustive match next to
  method()/params(): initialize HANDSHAKE_DEADLINE, ping PING_DEADLINE (10 s,
  also the heartbeat's default hang_timeout), shutdown 2 s (tests), activate
  ACTIVATE_DEADLINE + 1 s, deactivate DISPOSE_DEADLINE + 500 ms, tool/call
  its own deadline_ms, so the host is told exactly the bound the core
  enforces.
- HostProcess::call() replaces request() and request_with_deadline(): it
  applies the method's deadline, and on expiry or when the caller drops the
  future it sends `$/cancel` and forgets the call (CancelOnDrop moved here
  from tool.rs). HostCallError::Timeout now names the method.
- SupervisionOptions::tool_call_deadline (default tool::TOOL_CALL_DEADLINE,
  120 s) feeds ToolCallParams::deadline_ms; HostToolSpec::execute is one
  call() instead of its own guard and timeout.
- Only the heartbeat still drives start_request() directly, under its own
  ping_timeout/hang_timeout.

Test: extension_host::tests::
a_call_past_its_method_deadline_is_cancelled_and_the_host_stays_usable
(Node; skips without one unless CODEWHALE_EXT_HOST_TESTS is set). With a
300 ms tool deadline, slow_wait (30 s unless cancelled) returns
ToolError::Timeout in under 2 s; fixture_script_tool then succeeds on the same
pid with spawn_attempts == 1; disabling slow-tool tears down without a
"teardown" diagnostic, which proves the host received `$/cancel`
(slow-tool's disposer awaits the in-flight call and would outlive the 2 s
dispose deadline). Fix-off controls: make call() await rx unbounded (the
test's 5 s outer timeout fires); remove the cancel from CancelOnDrop::drop
(the teardown diagnostic appears).

Evidence: rustfmt --check clean; cargo not run here (the lead compiles).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
… request counts

- Conformance: the recording host's `## Environment` shell line is the
  dispatcher's `$SHELL` path verbatim (`/bin/zsh`, a Custom shell kind), not
  `zsh`. Batch 5 on macOS showed the replayed prefix hash drifting to a third
  value until the recorded value matched what the host actually rendered.
- Extension host: `requests_started` (test-only) now excludes the monitor's
  heartbeat pings. They share the request id counter and fire every 3s, so a
  ping inside the approval window made
  execute_tools_gates_an_extension_tool_before_any_host_call report a tool
  request that never happened. That window became reachable once Node, whose
  macOS host is slower to start, became the default runtime. Production
  counting is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Phase 0 (CURRENT_DECISIONS §26), visible "host down".

Before: HostStatus existed but only /plugin rendered it, with its own
strings; a tool call routed to a down host got "the extension host is not
running" (or, after a crash, "no longer registered", because liveness was
checked before the host), and ensure_host had four more ad-hoc strings
("unavailable; waiting for supervision", "is restarting", "is failed (…)",
"is starting").

Now HostStatus is the one source:
- `impl Display for HostStatus` is the one phrase per state: disabled by
  config, not started, starting, unresponsive (pid), restarting after:
  <exit reason>, or the failure reason plus "(change or reload a plugin to
  retry)". New variant HostStatus::Disabled for `[features] extension_host`
  off. Launch failures are recorded as "start failed: <reason>"; the crash
  budget keeps "crash budget exhausted: <reason>".
- HostSlot::status() derives it (moved from ExtensionHostManager::status); a
  Ready/Unresponsive slot whose process already exited reads as restarting
  instead of "running".
- ManagerShared::ready_host() returns Result<_, HostStatus>. live_host_for
  checks the policy and the host first, so a call to a down host fails with
  NotAvailable "extension host is down: <why>"; ensure_host uses the same
  host_down() text.
- render_status prints the Display phrase for every state but Ready's detail
  line (Failed keeps its stderr tail). status_report() returns a String and
  is one "disabled by config" line when the flag is off, so /plugin always
  says where the host stands.
- docs/EXTENSIONS.md states what users see and the deadline behaviour.

Test: extension_host::tests::
a_host_that_is_down_names_why_in_plugin_status_and_tool_errors (no Node
needed). For a crashed/restarting slot, a start-refused Failed slot and Idle,
render_status and HostToolSpec::execute on a live registration both carry
the same phrase ("restarting after: exited with signal: 9 (SIGKILL)",
"start failed: … initialization refused (change or reload a plugin to
retry)", "not started"); with the policy off, status_report() and the call
both say "disabled by config". Fix-off control: remove the early
`self.ready_host()` check in live_host_for (the call falls through to the
authority check and the message no longer names the host state).

Known limit: the phrases are English, like the rest of render_status; they
are not yet routed through tr().

Evidence: rustfmt --check clean; cargo not run here (the lead compiles).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
- Reassigning a `&Ty` from a `&Box<Ty>` inside `while let` leaned on
  assignment coercion; `innermost` takes the item type through a function
  argument instead.
- `kinds` takes the field slice and a required/optional flag instead of an
  `impl Iterator<Item = &'a Field>`, so no explicit lifetime is needed.

No change to the rendered TypeScript.

Evidence: rustfmt --check clean; cargo not run here (the lead compiles).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…d with

Instrumenting the harness showed the recording environment (the governed test
process) renders `shell="/bin/bash"` on macOS; `/bin/zsh` came from an
interactive shell, not the recorder. With the observed value the replayed
environment block matches the recording. Hosted Linux runners also report
/bin/bash, so there only the OS line differs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
The shell tool description and command guidance are process-wide OnceLock
caches built from the global dispatcher, so a shared-process test run
rendered whichever shell the first test to touch them saw (/bin/zsh on a
developer Mac) while the conformance sandbox replays /bin/bash. Parallel
runs of the event goldens then drifted in the tool catalog.

Under a pinned recorded environment the guidance and description render
from the recorded shell instead of the cache (test builds, unix only);
production keeps the cached dispatcher shell unchanged.

Refs #6698

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
The auth-required classifier matched the bare substring "401". Transport
errors carry the URL they failed on, so the connection reset in
streamable_http_reset_after_tool_call_post_is_not_replayed read as a 401
whenever the loopback port contained it (#6604's macOS run: port 50401,
twice). On a live call that misclassification drops the connection, flags
the server ◆ auth required and names an OAuth login the server never asked
for — for any real server whose URL or port carries those digits.

text_names_http_status matches a whole digit run outside URL tokens. The
doctor's 401/403 hints use it too. Known limit: word signals such as
"unauthorized" are still matched anywhere in the text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Generated host protocol TypeScript, the core-only authority lint,
per-method cancel deadlines and a visible host-down reason, reviewed in the
extension-host lane.

Signed-off-by: CodeWhale Bot <bot@codewhale.net>
);
}
let malformed = client
.delete(format!("http://{addr}/v1/sessions/not.a.session"))

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants