Skip to content

feat(voice): a call gets the real system prompt, so it can be delegated to - #2406

Merged
2witstudios merged 12 commits into
masterfrom
pu/realtime-prompt
Aug 12, 2026
Merged

2witstudios merged 12 commits into
masterfrom
pu/realtime-prompt

Conversation

@2witstudios

@2witstudios 2witstudios commented Aug 12, 2026

Copy link
Copy Markdown
Owner

A voice call could hold a conversation and reach about ten tools. It could not be delegated to, and the reason was not tone.

The bug

buildRealtimeToolSet hand-rolled the search-mode exposure split — splitToolsForExposure plus the two scaffolding factories — instead of calling applyToolExposureMode. That function returns the tool set and the discovery prompt naming every deferred tool, as one expression, precisely so the advertised tools and the text describing them cannot disagree. Voice reproduced the tool half and dropped the text half.

So tool_search and execute_tool rode every session while nothing in the instructions ever named them, and everything outside CORE_TOOL_NAMEScreate_task, spawn_session, the calendar family, the workflow tools — was loaded and undiscoverable. The model answered "I can't do that" about tools it was holding.

The prompt around them was thin for the same underlying reason: voice built its own five-bullet string rather than the assembly the typed surface uses, so a call also carried no workspace knowledge, no skill catalog, no agent memory and no plan pointer.

What this does

  • buildAgentSystemPrompt (core/prompt-assembly.ts) — the stable system prefix, extracted from page-chat-turn.ts and global-chat-turn.ts with both surfaces side by side. The order, and the blank-slate branch for an agent carrying its own prompt, lived in no single place and drifted; global-chat-turn.ts's own docblock records where that ended up ("it claimed tasks create linked DOCUMENT pages; they create TASK_LIST children"). Text output is unchanged — characterization tests reproduce the old expression by hand, and the moved Global Assistant literal was diffed byte-for-byte against HEAD.
  • realtime/system-context.ts gathers that assembly's inputs for whichever surface a call is bound to, and caps it with the spoken override. Every read is individually best-effort and names itself in the log, so a dead plan pointer costs the pointer, not the call.
  • realtime/instructions.ts is now only what changes because the words are heard, appended last as an explicit override block. gpt-realtime degrades on conflicting instructions specifically, so conflicts are named and resolved ("Skip preambles" does NOT apply here) rather than left for the model to arbitrate.
  • realtime/tools.ts calls applyToolExposureMode and carries the whole exposure out: the tool set, both halves of the discovery text, and the pre-split capability names.
  • core/complete-request-builder.ts — the admin "exact context window" viewer was the third copy, and had already drifted: it omitted the Global Assistant's exploration guidance entirely and called the capability builders with the no-filtering sentinel. It now calls the shared builder and honours its own contextType, pinned by a byte-identity test.

A call now also stops asking permission before acting, hands multi-minute work to spawn_session / create_task instead of leaving the caller in silence, and knows that "this page" resolves without an id.

Review round 1 (aa5e724)

Six findings from chatgpt-codex-connector and coderabbitai, all valid:

  • Capability names were read after the split, not before it. Object.keys(exposure.tools) is the post-split set, so every deferred capability looked disabled: buildInlineInstructions dropped TASK MANAGEMENT / AGENTS / AUTOMATION, and the skill catalog dropped task-management and spreadsheets, while writing-documents survived on core tools and kept the section looking populated. The exposure now returns allowedToolNames itself, captured where page-chat-turn.ts:1306 captures it.
  • tool_search had no skills to search. searchableSkills was a parameter nobody passed. Now computed inside the exposure from the pre-split names; the parameter is gone.
  • Rules named tools that were not there. A core-only agent is registered no scaffolding by design, yet the override still ordered it to call tool_search, and still offered spawn_session/create_task. Both are now conditioned on the exposed and reachable sets. The shared builder had the same problem on the global surface and is gated on the catalog it introduces.
  • A whole-deployment scan on the handshake. buildAgentAwarenessPrompt selects every non-trashed drive, then awaits an access check per drive and a view check per agent serially — before the SDP exchange. Dropped from the call path; list_agents fetches the list on the turn that needs it. The query's own inefficiency is pre-existing on the typed surface and left alone.
  • Delegation lost a clause. Making the hand-offs conditional dropped "a trigger or a workflow for work that should happen later" — found by reading the generated prompt end to end, not by a test. Restored, keyed on either set_task_trigger or create_workflow.
  • The plan pointer is a directive that cannot be re-sent. A clear_plan mid-call would leave the model told to resume a plan the caller just ended. A session.update path is the real fix and is out of scope; the override now names the model's own tool result as authoritative for anything it changes during the call, which generalises to agent memory too. Documented as a mitigation, not a mechanism.

Follow-on hardening

The exposure was being computed twice per call — once in route.ts for the advertised tool definitions, once in system-context.ts for the prompt describing them. Deterministic, so they agreed; also exactly the arrangement that produced the original bug. buildVoiceCallContext now returns { instructions, tools } from one exposure, the binding carries both, and the route forwards what it was given. That also takes a second full registry build off the handshake the caller is waiting through.

Pinned by a test asserting every advertised tool name appears in the prompt shipped with it, and its converse for a core-only agent.

Two behavior changes beyond the refactor, both deliberate

  • A whitespace-only custom systemPrompt is no longer a prompt. It was truthy on the typed surface too, so one stray space in the field suppressed the default persona and the workspace knowledge, leaving an agent whose entire brief was " ".
  • An agent allowed only core tools no longer gets tool_search/execute_tool over an empty catalog. That is applyToolExposureMode's own rule, now shared instead of re-decided: two tools whose every call fails are worse than two tools absent.

Deliberately still omitted from a call

The page tree, the drive prompt, and the cross-drive member context. Instructions ride a single session.update at socket open and there is no path that sends a second, so a drive's instructions frozen at connect time would go wrong the moment the caller walked to another drive. The tools read the live location instead. Reasons are recorded in system-context.ts.

Verification

  • Monorepo typecheck (incl. build + lint), lint, knip green.
  • 17,249 web tests pass. The 18 failing files all require Postgres and fail on the connection.
  • CI: Security Test Suite, Static Security Analysis, CodeQL, Dependency Audit, Secret Scanning, Lint & TypeScript Check and the agent-session E2E suite all green.
  • The page/global duplication ratchet moved down to 164 — lowered in all three recorded homes rather than raised.
  • Mutation-checked throughout. Five of my own tests were caught being wrong across the two rounds — passing for the wrong reason (matching tool names elsewhere in the prose; asserting guidance on the branch that states it unconditionally), or asserting something untrue (a read-only preview dropping create_task, which the global surface never gated). Each was tightened until the mutation went red, or corrected.
  • Assembled prompt measures ~13k chars global / ~11k page ≈ 3,200 tokens against a 32k session; pinned by a ceiling test.

Not verified: a real call. Everything here is unit-level, and a prompt is only really tested by talking to it. The check that matters is "What's on my calendar tomorrow?" → expect tool_searchexecute_tool rather than a refusal.

If it over-fires — calling tools before the caller finishes a sentence — the remedy is to soften rather than delete: lower-case the capitalized DO NOT ASK PERMISSION and leave the rest standing.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd

Summary by CodeRabbit

  • New Features

    • Voice calls now retain the assistant’s instructions, workspace knowledge, personalization, plans, and available tools.
    • Assistants can perform requested actions during calls without separate permission prompts, delegate longer tasks, and report outcomes.
    • Voice interactions better understand references such as “this page” using the current view.
    • Tool availability and discovery are now tailored to each conversation.
  • Documentation

    • Added an unreleased changelog entry describing voice-call delegation capabilities.

…ed to

A voice call could hold a conversation and reach about ten tools. It could not
be delegated to, and the reason was not tone.

`buildRealtimeToolSet` hand-rolled the search-mode exposure split
(`splitToolsForExposure` + the two scaffolding factories) instead of calling
`applyToolExposureMode`, and so kept only half of what that function returns.
The other half is the discovery prompt naming every deferred tool. They are one
expression there precisely so the advertised tools and the text describing them
cannot disagree — and voice dropped the text. `tool_search` and `execute_tool`
rode every session while nothing in the instructions ever named them, so
`create_task`, `spawn_session`, the calendar family and the workflow tools were
loaded and undiscoverable. The model answered "I can't do that" about tools it
was holding.

The prompt around them was thin for the same reason nobody noticed: voice built
its own five-bullet string rather than the assembly the typed surface uses, so
it also carried no workspace knowledge, no skill catalog, no agent memory and no
plan pointer.

WHAT THIS DOES

- `buildAgentSystemPrompt` (core/prompt-assembly.ts) — the stable system prefix,
  extracted from page-chat-turn.ts and global-chat-turn.ts with both surfaces
  side by side. The order and the blank-slate branch are the parts that lived in
  no single place and drifted; global-chat-turn's own docblock records where
  that ended up ("it claimed tasks create linked DOCUMENT pages").
- realtime/system-context.ts gathers that assembly's inputs for whichever
  surface a call is bound to. Every read is individually best-effort and names
  itself in the log: a dead plan pointer costs the pointer, not the call.
- realtime/instructions.ts is now only what changes because the words are HEARD,
  appended last as an explicit override block. gpt-realtime degrades on
  conflicting instructions specifically, so the conflicts are named and resolved
  ("Skip preambles" does NOT apply here) rather than left to the model.
- realtime/tools.ts calls applyToolExposureMode and carries both halves out.

TWO BEHAVIOR CHANGES BEYOND THE REFACTOR, BOTH DELIBERATE

- A whitespace-only custom systemPrompt is no longer a prompt. It was truthy on
  the typed surface too, so one stray space in the field suppressed the default
  persona AND the workspace knowledge, leaving an agent whose entire brief was
  "   ".
- An agent allowed only core tools no longer gets tool_search/execute_tool over
  an empty catalog. That is applyToolExposureMode's own rule, now shared instead
  of re-decided: two tools whose every call fails are worse than two tools
  absent.

The page/global duplication ratchet moved DOWN to 164 and is lowered in all
three recorded homes. Deliberately still omitted from a call: the page tree, the
drive prompt and the cross-drive member context — instructions ride a single
session.update at socket open and there is no path that sends a second, so a
drive's instructions frozen at connect time would go wrong the moment the caller
walked to another drive. The tools read the live location instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
@coderabbitai

coderabbitai Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 497ff220-ff81-4046-b8f9-5da8fe2a1e92

📥 Commits

Reviewing files that changed from the base of the PR and between 1b6818f and 2b2d819.

📒 Files selected for processing (9)
  • apps/web/src/app/api/voice/realtime/call/__tests__/route.test.ts
  • apps/web/src/app/api/voice/realtime/call/route.ts
  • apps/web/src/lib/ai/core/__tests__/complete-request-builder.test.ts
  • apps/web/src/lib/ai/core/complete-request-builder.ts
  • apps/web/src/lib/ai/realtime/__tests__/binding-loader.test.ts
  • apps/web/src/lib/ai/realtime/__tests__/system-context.test.ts
  • apps/web/src/lib/ai/realtime/binding-loader.ts
  • apps/web/src/lib/ai/realtime/system-context.ts
  • apps/web/src/lib/ai/realtime/voice-runtime-deps.ts
🚧 Files skipped from review as they are similar to previous changes (3)
  • apps/web/src/lib/ai/realtime/binding-loader.ts
  • apps/web/src/lib/ai/realtime/tests/binding-loader.test.ts
  • apps/web/src/lib/ai/realtime/tests/system-context.test.ts

📝 Walkthrough

Walkthrough

The PR centralizes page and global prompt assembly, extends the shared path to realtime voice calls, applies request-scoped tool exposure and context loading, and adds coverage for prompt content, voice behavior, authorization paths, failures, and stable output.

Changes

Unified prompt assembly

Layer / File(s) Summary
Shared prompt assembly and chat integration
apps/web/src/lib/ai/core/..., apps/web/src/lib/ai/chat-pipeline/..., docs/2.0-architecture/agent-sessions.md
buildAgentSystemPrompt now assembles page-agent and Global Assistant prompts. Chat turns pass stable prompt inputs without repeated concatenation.
Realtime context, tools, and voice instructions
apps/web/src/lib/ai/realtime/..., apps/web/src/lib/ai/tools/tool-exposure.ts
Realtime calls build shared agent context, apply allowlisted tool exposure, and append voice-specific execution instructions.
Authenticated binding integration
apps/web/src/lib/ai/realtime/voice-runtime-deps.ts, apps/web/src/lib/ai/realtime/binding-loader.ts, apps/web/src/app/api/voice/realtime/call/..., CHANGELOG.md
Voice dependencies use authenticated request factories. Bindings receive generated instructions and tools. The changelog documents voice delegation behavior.

Estimated code review effort: 4 (Complex) | ~45 minutes

Mergeability Score: ⚪ Minimal · up to 2b2d8

The change gives voice calls the shared system prompt and matching tool discovery/exposure while removing an unnecessary pre-call scan; no actionable merge-blocking risk remains after normal checks and review.

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant VoiceCallRoute
  participant loadVoiceBinding
  participant buildVoiceCallContext
  participant buildAgentSystemPrompt
  participant VoiceRuntime
  Caller->>VoiceCallRoute: authenticated realtime call
  VoiceCallRoute->>loadVoiceBinding: request-scoped dependencies
  loadVoiceBinding->>buildVoiceCallContext: conversation and agent metadata
  buildVoiceCallContext->>buildAgentSystemPrompt: tools, memory, plans, and personalization
  buildAgentSystemPrompt-->>buildVoiceCallContext: stable agent system prompt
  buildVoiceCallContext-->>loadVoiceBinding: voice instructions and realtime tools
  loadVoiceBinding-->>VoiceRuntime: voice binding
  VoiceRuntime-->>Caller: spoken response or delegated action result
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 60.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: voice calls receive the real system prompt and support delegation.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch pu/realtime-prompt

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e3e009aac2

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

): Promise<string> => {
const allowlist: ToolAllowlist = request.agent?.enabledTools ?? null;
const exposure = buildRealtimeToolExposure(deps.buildTools(), allowlist);
const allowedToolNames = Object.keys(exposure.tools);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Derive prompt capability names before search exposure

When an enabled capability is deferred, such as create_task, spawn_session, or edit_sheet_cells, exposure.tools contains only core tools plus the search scaffolding. Using those keys as allowedToolNames makes buildInlineInstructions and buildBuiltinSkillCatalog treat the deferred capabilities as disabled, so page agents lose task, delegation, and automation guidance and omit skills such as task management and spreadsheets even though the non-core catalog says those tools are callable. Derive this list from the allowlist-filtered registry before applying search exposure.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and this was the sharpest finding on the PR — thank you. Fixed in aa5e724.

Object.keys(exposure.tools) is the post-split set, so every deferred capability read as disabled: buildInlineInstructions dropped TASK MANAGEMENT, AGENTS and AUTOMATION, and buildBuiltinSkillCatalog dropped task-management and spreadsheets (their requiredTools are create_task/update_task and edit_sheet_cells). writing-documents survived on core tools alone and kept the SKILLS section looking populated, which is what made it easy to miss.

Rather than deriving the list at the call site, buildRealtimeToolExposure now returns allowedToolNames itself — captured after the allowlist filter and before the split, the same point page-chat-turn.ts:1306 captures it. The wrong list is no longer reachable from the caller.

Pinned by a test that asserts the guidance sections on the page branch specifically (the Global Assistant's builder states them unconditionally and would have passed either way), and mutation-checked: restoring Object.keys(exposure.tools) turns it red.

Comment on lines +124 to +128
request.conversationId
? softly(
deps,
'activePlan',
() => deps.loadActivePlan(request.conversationId as string, request.userId),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Refresh active-plan instructions after plan mutations

For a call that invokes set_plan or clear_plan, this lookup freezes the pre-call plan binding into session instructions that are never updated. Subsequent utterances in the same call therefore either lack the newly selected plan or continue receiving the old system-level ACTIVE PLAN directive after it was cleared, which can make the assistant re-read and continue working against a stale plan. Mutable plan state needs a session.update path or must be obtained dynamically instead of being captured only at handshake.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Real, and specific to this surface — thank you. Addressed in aa5e724, though not the way you suggest, and I want to be explicit about the tradeoff.

The ACTIVE PLAN block is a directive rather than a fact ("keep working against this page, re-read it before continuing, call clear_plan when finished"), so a stale one is actively wrong and not merely missing. You are right that a clear_plan mid-call leaves the model being told to resume a plan the caller just ended.

A session.update path is the correct fix and is genuinely out of scope here: it needs a new contract message, a sender in apps/realtime, and care around the empty-instructions footgun at session.ts:72-90, where an empty instructions replaces the session persona rather than leaving it alone.

What I did instead uses the one channel that can already correct the record mid-call — the model's own tool result:

Everything above was assembled when this call started and is never re-sent. If you change any of it during the call — bind or clear a plan, edit your memory page, move or rename something — YOUR OWN TOOL RESULT IS WHAT IS CURRENT... Never re-follow a standing instruction about something you have since changed.

That generalises past the plan to agent memory and anything else in the same class. It is a mitigation, not a mechanism, and I would rather say so than claim it is equivalent. If you think the residual risk warrants blocking on the real session.update path, say so and I will scope it.

const agentAwareness = await softly(
deps,
'agentAwareness',
() => deps.loadAgentAwareness(request.userId),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Avoid scanning every drive during the voice handshake

On every unbound or global call, this read blocks before runCallHandshake performs the SDP exchange. buildAgentAwarenessPrompt first selects every non-trashed drive in the deployment, then awaits getUserDriveAccess serially for each one and performs additional sequential agent permission checks, so voice setup latency and database work grow with all tenants' drives rather than this user's drives. On a populated installation this can substantially delay call establishment; use a membership-scoped query or move this work off the handshake's critical path.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confirmed by reading it, and worse than "grows with drives" — buildAgentAwarenessPrompt selects every non-trashed drive in the deployment, then awaits getUserDriveAccess per drive serially, then canUserViewPage per agent serially. All of it before the SDP exchange, with the caller holding a dead line waiting to be heard.

Dropped from the call path in aa5e724. The capability is not lost: the AGENTS guidance is still in the prompt, and list_agents / multi_drive_list_agents fetch the list on the one turn that actually needs it rather than on every call that might — which is the better trade on a surface where the standing prompt cannot be re-sent anyway.

I have deliberately not rewritten the query itself. It is inefficient on the typed surface too, but that is a shared function on a non-latency-critical path and it deserves its own change rather than riding along here. Flagging it as worth a follow-up.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (1)
apps/web/src/lib/ai/realtime/__tests__/tools.test.ts (1)

262-324: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Add coverage for the new nonCoreToolNames field.

This block pins toolDiscoveryPrompt well, including the registry-wide invariant at Lines 307-323. RealtimeToolExposure also exposes nonCoreToolNames, and system-context.ts sends that value to the global surface as the standalone catalog. No test asserts it.

Pin two properties: the catalog names the deferred tools, and it is '' when nothing is deferred.

♻️ Suggested additional cases
   it('given nothing to defer, should return no prompt rather than an empty instruction', () => {
     expect(buildRealtimeToolExposure({ read_page: fakeTool() }).toolDiscoveryPrompt).toBe('');
     expect(buildRealtimeToolExposure({}).toolDiscoveryPrompt).toBe('');
   });
+
+  it('should return the catalog alone for the surface that states the "how" earlier', () => {
+    // The global surface takes nonCoreToolNames on its own and states
+    // TOOL_DISCOVERY_PROMPT itself, so the two halves must stay consistent.
+    const { nonCoreToolNames } = buildRealtimeToolExposure(smallSet());
+
+    expect(nonCoreToolNames).toContain('rename_drive');
+    expect(nonCoreToolNames).not.toContain('read_page');
+  });
+
+  it('given nothing to defer, should return an empty catalog', () => {
+    expect(buildRealtimeToolExposure({ read_page: fakeTool() }).nonCoreToolNames).toBe('');
+    expect(buildRealtimeToolExposure({}).nonCoreToolNames).toBe('');
+  });
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/web/src/lib/ai/realtime/__tests__/tools.test.ts` around lines 262 - 324,
Add tests in the “buildRealtimeToolExposure — the discovery prompt” suite for
the `nonCoreToolNames` field: verify it includes every deferred tool name from a
registry, and verify it equals `''` when the registry has no deferred tools. Use
the existing `smallSet()` and `buildPageSpaceTools`/allowlist scenarios where
appropriate, without changing the existing prompt assertions.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@apps/web/src/lib/ai/realtime/instructions.ts`:
- Around line 85-86: The voice instructions unconditionally require tool_search
even for core-only agents that do not expose discovery tools. In
apps/web/src/lib/ai/realtime/instructions.ts:85-86, pass discovery availability
into buildVoiceInstructions and omit or replace that rule when tool_search is
unavailable; update
apps/web/src/lib/ai/realtime/__tests__/instructions.test.ts:72-76 to assert the
rule only with discovery tools, and add a core-only registry case in
apps/web/src/lib/ai/realtime/__tests__/system-context.test.ts:75-105 verifying
tool_search, execute_tool, and their guidance are absent.

In `@apps/web/src/lib/ai/realtime/system-context.ts`:
- Around line 166-184: The global realtime prompt currently invokes
buildAgentSystemPrompt with an empty deferred catalog while the builder still
emits discovery instructions. Update the global branch around
buildAgentSystemPrompt, or the shared builder’s global-surface logic, to include
discovery text only when exposure.nonCoreToolNames is non-empty, preserving the
existing prompt when deferred tools are available.
- Around line 118-121: Update the realtime context construction around
buildRealtimeToolExposure and allowedToolNames to preserve the pre-exposure
names from deps.buildTools() for buildBuiltinSkillCatalog and the inline/global
instruction builders, while retaining the exposed names for actual exposure
behavior. Build the eligible skill catalog before exposure and pass it as
searchableSkills to buildRealtimeToolExposure so tool_search can resolve
advertised skills.

---

Nitpick comments:
In `@apps/web/src/lib/ai/realtime/__tests__/tools.test.ts`:
- Around line 262-324: Add tests in the “buildRealtimeToolExposure — the
discovery prompt” suite for the `nonCoreToolNames` field: verify it includes
every deferred tool name from a registry, and verify it equals `''` when the
registry has no deferred tools. Use the existing `smallSet()` and
`buildPageSpaceTools`/allowlist scenarios where appropriate, without changing
the existing prompt assertions.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 1a342ec5-48ad-4347-b3b7-d11318529564

📥 Commits

Reviewing files that changed from the base of the PR and between 2ffd879 and e3e009a.

📒 Files selected for processing (21)
  • CHANGELOG.md
  • apps/web/src/app/api/voice/realtime/call/__tests__/route.test.ts
  • apps/web/src/app/api/voice/realtime/call/route.ts
  • apps/web/src/lib/ai/chat-pipeline/__tests__/turn-duplication-ratchet.test.ts
  • apps/web/src/lib/ai/chat-pipeline/global-chat-turn.ts
  • apps/web/src/lib/ai/chat-pipeline/handle-chat-turn.ts
  • apps/web/src/lib/ai/chat-pipeline/page-chat-turn.ts
  • apps/web/src/lib/ai/core/__tests__/agent-system-prompt.test.ts
  • apps/web/src/lib/ai/core/prompt-assembly.ts
  • apps/web/src/lib/ai/core/system-prompt.ts
  • apps/web/src/lib/ai/realtime/__tests__/binding-loader.test.ts
  • apps/web/src/lib/ai/realtime/__tests__/instructions.test.ts
  • apps/web/src/lib/ai/realtime/__tests__/system-context.test.ts
  • apps/web/src/lib/ai/realtime/__tests__/tools.test.ts
  • apps/web/src/lib/ai/realtime/binding-loader.ts
  • apps/web/src/lib/ai/realtime/instructions.ts
  • apps/web/src/lib/ai/realtime/system-context.ts
  • apps/web/src/lib/ai/realtime/tools.ts
  • apps/web/src/lib/ai/realtime/voice-runtime-deps.ts
  • apps/web/src/lib/ai/tools/tool-exposure.ts
  • docs/2.0-architecture/agent-sessions.md

Comment thread apps/web/src/lib/ai/realtime/instructions.ts Outdated
Comment thread apps/web/src/lib/ai/realtime/system-context.ts
Comment thread apps/web/src/lib/ai/realtime/system-context.ts Outdated
Six review findings, all of them real. The first two are the same bug and it
undercut the change this PR exists to make.

CAPABILITY NAMES WERE READ AFTER THE SPLIT, NOT BEFORE IT

`allowedToolNames` came from `Object.keys(exposure.tools)`. After the exposure
split that object holds the core tools and the two scaffolding tools and nothing
else, so every capability the split deferred looked disabled to the things gated
on that list. `buildInlineInstructions` dropped TASK MANAGEMENT, AGENTS and
AUTOMATION; `buildBuiltinSkillCatalog` dropped `task-management` and
`spreadsheets` while `writing-documents` survived and kept the section looking
populated. The catalog went on advertising create_task and spawn_session as
callable with every word of guidance about them removed — the model told the
tools exist and nothing about when to use them.

`page-chat-turn.ts:1306` captures the same list before exposure with a comment
saying why. The exposure now returns `allowedToolNames` itself, so the wrong
list is no longer reachable from the call site, and returns `eligibleSkills`
with it — computed internally rather than accepted as a parameter no caller was
passing, which is why `tool_search` could not resolve a skill the prompt
advertised.

RULES THAT NAMED TOOLS THAT WERE NOT THERE

An agent allowed only core tools is registered no scaffolding, by design. The
override still ordered it to call `tool_search` before refusing anything, and
still offered `spawn_session` and `create_task` as hand-offs. Both are now
conditioned on the exposed and reachable sets respectively, which is why the
override needs `VoiceToolReach` and why `buildVoiceSystemContext` — the only
place holding both halves — now returns the finished instructions rather than
just the system prompt. The shared builder had the same problem on the global
surface, stating TOOL_DISCOVERY_PROMPT with nothing deferred; gated on the
catalog it introduces, which is unreachable for the text routes.

A WHOLE-DEPLOYMENT SCAN ON THE HANDSHAKE

`buildAgentAwarenessPrompt` selects every non-trashed drive in the deployment,
then awaits an access check per drive and a view check per agent, one after
another — all of it before the SDP exchange, with the caller holding a dead line
waiting to be heard. The eager list is dropped from the call path. The AGENTS
guidance stays, and `list_agents` fetches the list on the one turn that needs it
rather than on every call that might. The query itself is inefficient on the
typed surface too; that is not this PR's to fix.

INSTRUCTIONS THAT CANNOT BE RE-SENT

A call's instructions are assembled once at socket open, so the ACTIVE PLAN
pointer — a directive, not a fact ("keep working against this page, re-read it
before continuing") — would go on being followed after a `clear_plan` on the
same call. The one channel that can correct the record mid-call is the model's
own tool result, so the override now names it as authoritative for anything the
model changes while talking. That covers the plan, the memory page, and
everything else in the same class.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
apps/web/src/lib/ai/realtime/binding-loader.ts (1)

103-108: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick win

Make the fallback independent of buildInstructions.

unbound awaits deps.buildInstructions at Line 108. If that promise rejects, loadVoiceBinding catches it at Lines 207-215 and calls unbound again. The same promise can reject again, so the call fails instead of falling back.

Catch instruction-building failures once and return a non-throwing baseline instruction value. Do not route this fallback through deps.buildInstructions again.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/web/src/lib/ai/realtime/binding-loader.ts` around lines 103 - 108,
Update unbound so its fallback does not call or await deps.buildInstructions;
return the baseline non-throwing instruction value directly while preserving the
empty seed. Ensure the loadVoiceBinding recovery path can invoke unbound without
repeating the failing instruction-building operation.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@apps/web/src/lib/ai/realtime/binding-loader.ts`:
- Around line 103-108: Update unbound so its fallback does not call or await
deps.buildInstructions; return the baseline non-throwing instruction value
directly while preserving the empty seed. Ensure the loadVoiceBinding recovery
path can invoke unbound without repeating the failing instruction-building
operation.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: c383f35c-09ff-4c55-8a79-d428d6c4e5e6

📥 Commits

Reviewing files that changed from the base of the PR and between e3e009a and aa5e724.

📒 Files selected for processing (10)
  • apps/web/src/lib/ai/core/__tests__/agent-system-prompt.test.ts
  • apps/web/src/lib/ai/core/prompt-assembly.ts
  • apps/web/src/lib/ai/realtime/__tests__/binding-loader.test.ts
  • apps/web/src/lib/ai/realtime/__tests__/instructions.test.ts
  • apps/web/src/lib/ai/realtime/__tests__/system-context.test.ts
  • apps/web/src/lib/ai/realtime/binding-loader.ts
  • apps/web/src/lib/ai/realtime/instructions.ts
  • apps/web/src/lib/ai/realtime/system-context.ts
  • apps/web/src/lib/ai/realtime/tools.ts
  • apps/web/src/lib/ai/realtime/voice-runtime-deps.ts
🚧 Files skipped from review as they are similar to previous changes (6)
  • apps/web/src/lib/ai/core/tests/agent-system-prompt.test.ts
  • apps/web/src/lib/ai/realtime/tests/binding-loader.test.ts
  • apps/web/src/lib/ai/realtime/tests/instructions.test.ts
  • apps/web/src/lib/ai/realtime/tests/system-context.test.ts
  • apps/web/src/lib/ai/realtime/voice-runtime-deps.ts
  • apps/web/src/lib/ai/core/prompt-assembly.ts

2witstudios and others added 10 commits August 12, 2026 16:20
…shape

Three cleanups on the review fixes, all found by rereading the diff as a
reviewer would:

- `RealtimeToolExposure.eligibleSkills` was returned and never read. The value
  is real — it is the corpus `tool_search` gets — but it is consumed inside the
  exposure, so exporting it was documentation pretending to be an interface.
- `buildVoiceSystemContext` built its argument object inline inside the call
  that wrapped it, which read as one expression doing two things. Each branch
  now names its system prompt and then caps it.
- `HAND_OFFS` was rebuilt on every call and sat between a docblock and the
  function it documented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
Making delegation conditional on reachable tools dropped a clause the original
override had: work that should happen LATER, via a trigger or a workflow. A
call is exactly where that matters — the caller cannot sit and wait, so
"later" is often the right answer and the model had stopped being told it was
available.

Restored as a third entry, keyed on either set_task_trigger or create_workflow
rather than one tool: the phrase names a destination, not a call, so an agent
that can set a trigger but not build a workflow should still hear it.

Caught by reading the generated prompt end to end rather than by a test — worth
noting, since every existing test asserts fragments and none of them would have
missed a whole missing clause.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
The last piece of the extraction. `complete-request-builder.ts` renders a page
titled "the exact context window", and it was assembling its own prompt — the
third copy, and one that had already drifted:

- it omitted the Global Assistant's exploration guidance ENTIRELY, so an admin
  reading it never saw the rules that decide which drive "here" means;
- it called the capability builders with no arguments, which is the
  no-filtering-context sentinel, so the sections it showed belonged to no
  caller in particular;
- it ordered the blocks in a way neither route used.

It now calls `buildAgentSystemPrompt`, and honours its own contextType instead
of always previewing the global surface. The blocks it genuinely cannot know —
plan pointer, agent memory, drive context — are passed empty with the reason
stated, rather than being silently absent.

Pinned by a byte-identity test against the shared builder, because this file's
whole claim is exactness and "looks about right" is the failure mode.

One test of my own was wrong before it was right: I asserted a read-only
preview would drop `create_task`, which the global surface never gated. Fixed
to assert what the change actually does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
The route built the advertised tool definitions and `system-context` built the
prompt describing them, each calling the exposure separately with the same
inputs. Deterministic, so they agree — today. It is also the precise
arrangement that produced the bug this PR exists to fix: a session advertising
`tool_search` while the prompt never named it.

They are two projections of one decision (what this agent may reach), so the
exposure is computed once and both come out of it. `buildVoiceCallContext`
returns `{ instructions, tools }`, the binding carries them together, and the
route forwards what it was given instead of rebuilding.

Also removes a second full registry build and exposure split from the
handshake — the path the caller is waiting through, and the one codex flagged
for latency.

Pinned by a test that every advertised tool name appears in the prompt that
ships with it, and its converse for a core-only agent: no scaffolding
advertised, none promised.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
`buildVoiceCallContext` returns the tools alongside the instructions, and the
header still said the module produced a system prompt. Also states the reason
the two travel together, which is the whole reason the module is shaped this
way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
The context ceiling was measured with nothing loaded — the case least likely to
breach it. The realistic heavy call is a bound agent carrying a memory page at
its own ~2k-token cap, a plan pointer, and personalization the user wrote, and
that is what has to fit beside 4k of seed and the audio for the length of the
call.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
Consolidating the exposure left `buildRealtimeTools` with no production caller
— only tests kept it compiling, which is exactly the kind of dead export knip
cannot see, since a test importing it counts as a use.

Replaced with `toRealtimeTools(toolSet)`: the projection both sides actually
need, taking the already-exposed set rather than a registry and an allowlist.
Re-deriving the exposure inside it is what would put the advertised list and
the prompt describing it back on separate computations, which is the thing this
PR spent two rounds removing.

The tests keep every case — schema conversion guard, flat wire shape, allowlist
behaviour — through a two-line local helper, so coverage is unchanged and only
the seam moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
…ing'

The entry claimed a call carries "the same instructions and the same knowledge
of your workspace" as the typed surface. It carries the same OPERATING
knowledge — how tasks, agents, automations and search work, the skill catalog,
the tools behind the discovery catalog — but deliberately not the page tree,
the drive prompt or the cross-drive member context, each omitted for a reason
recorded in system-context.ts. Claiming parity we do not have is the kind of
thing a changelog should not do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DEofnKqGYY8Gs8CjwkGPXd
@2witstudios
2witstudios merged commit 64aeb4f into master Aug 12, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant