Skip to content

feat(cookbooks): store threads in Gateway conversations and run tools with runTools() - #1258

Merged
vishxrad merged 31 commits into
mainfrom
visharad/cookbooks-gateway-conversations
Sep 30, 2026
Merged

vishxrad merged 31 commits into
mainfrom
visharad/cookbooks-gateway-conversations

Conversation

@vishxrad

@vishxrad vishxrad commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

This PR now includes #1261 (runTools() and raw completion chunks), since both changed the same chat routes. #1264 (chat library) is stacked on top.

Summary

All three cookbooks (conversational analytics, document comparison, and booking assistant) now keep their threads in Gateway conversations and run their tools with the OpenAI SDK's runTools().

Thread storage

  • Browser storage. Agent Interface gets storage={useOpenuiCloudStorage({ token: "/api/frontend-token", features: { artifact: false } })}. The browser lists, creates, renames, and deletes threads directly with the Conversations API.
  • Frontend token route. A new /api/frontend-token route mints a short-lived token with THESYS_API_KEY. It uses DEMO_USER_ID (default demo-user) and APP_ID (default per cookbook, for example booking-assistant-cookbook), the same variables apps from openui create use. The key stays on the server, and each cookbook sees only its own threads even when all three share a key.
  • Turn storage. When each run ends, the chat route appends the turn to the thread's conversation with storeChatCompletionHistory() from @openuidev/server (new dependency, pinned to 0.1.0). The turn is the user's message, each tool call and its result, and the answer, collected from the runner's message events and stored once runner.done() settles. The conversation id is the threadId that Agent Interface's storage created. A stopped or failed run still stores the user's message and any tool steps that finished, so the question stays after a reload; a partly streamed answer isn't kept. A failed write is logged without affecting the answer. The response ends only after the write, so a follow-up finds the turn.
  • History from the conversation. The route uses only the new question from the browser and loads the earlier turns, tool calls and results included, from the thread's Gateway conversation. src/lib/gateway-history.ts pages through GET /v1/conversations/{id}/items and converts the items back to Chat Completions messages, the reverse of storeChatCompletionHistory(), dropping tool calls that have no result (a stopped run). The browser can't supply this history: it never receives tool results over Chat Completions, so its copy has tool calls without results, which Gateway rejects with a 400. This way the model sees where each earlier answer came from, and conversation() in the route is the place to add compaction later.

Generation

  • Tool loop removed. Each cookbook's src/lib/tool-loop.ts is gone. The chat route calls gateway.chat.completions.runTools() with the tool's JSON schema and its executor.
  • The runner's own stream. The route returns runner.toReadableStream(), which carries every completion's chunks as one JSON object per line (application/x-ndjson), and the page reads it with openAIReadableStreamAdapter() and openAIMessageFormat. .asResponse() isn't available here: it belongs to a single create() call, while runTools() makes one request per round and reads each response itself to find the tool calls.
  • Plain request handling. Like the other examples and the openui-cloud template, the routes read the body with request.json() and trust the local page. The chat route returns 400 without a threadId or a closing user message. The zod request schema, 1 MB body reader, origin checks, and 503s for a missing key or missing data are gone; a missing key or unprepared data now throws.
  • Same round caps and first-round errors. maxChatCompletions is 3, 4, and 3. The route waits for the first chunk before responding, so a rejected key, rate limit, or unknown model returns an HTTP error with Gateway's message. Gateway reports an unknown model inside a successful response, so waiting only for the response let it through as a broken stream.

Docs

  • Storage in one sentence. Each cookbook page covers thread storage in a one-sentence "Thread storage" callout that links storeChatCompletionHistory(), and each diagram shows the conversation. The pages don't describe the Conversations API or frontend tokens. The READMEs keep those details, including what to change before deploying: sign users in, mint tokens for the signed-in user, and check that a threadId belongs to that user before storing a turn in it.
  • Walkthrough. The pages call generated output OpenUI Lang, name runTools() only in code and References, and drop paragraphs that repeated earlier steps. Each page ends with one References list of the packages, APIs, and guides the example uses. Mentions of authentication link to the Gateway authentication page, and the booking prompt no longer calls search_stays a Responses tool.
  • Code samples fit the docs column. Sample lines are at most 80 characters and blocks at most 30 lines, so they don't scroll at desktop width. Document comparison's Agent Interface sample is split into its props and its slots. .prettierrc.json turns off embedded-code formatting for docs/content/docs/cookbooks/*.mdx, since Prettier would otherwise rewrap the samples at 100 characters. A few library.ts and sources.tsx sample lines here are still longer (up to 156 characters), because refactor(cookbooks): answer with the chat library #1264 replaces those samples.
  • Diagrams. Stage labels are unnumbered, since the numbered steps under each diagram follow a different sequence.

What changes for readers

  • No tool results in Behind the scenes during a run. Chat Completions has no chunk for a tool result, so Behind the scenes shows each call with its arguments but no result.
  • Parallel tools. Tools the model calls in one completion run in parallel, and results go back to the model in call order.
  • Tool errors. runTools() ends the run when a tool throws, so each route returns the error to the model as the tool result.
  • Later errors are generic. An error after the first chunk ends the stream, and Agent Interface shows "Something went wrong: Failed to fetch" instead of Gateway's message.
  • No forced final answer. A run that hits maxChatCompletions on a tool call ends without an answer.

Gateway items endpoint (fixed)

storeChatCompletionHistory() posts to POST /v1/conversations/{id}/items. Until 30 Sep 2026, Gateway answered valid requests with 404 "Conversation not found" for conversations that every other endpoint found. It now returns 200, the repro script prints "Not reproduced", and a reloaded thread loads its messages.

Test plan

  • npm ci and npm run verify (spec generation and next build with type checking) in all three cookbooks at this PR's head, rebased on main
  • In the browser, all three: the token route returns 200, GET /v1/conversations returns 200, and the first message creates a conversation (201) and streams the answer into a new thread
  • With toReadableStream(), in the browser: analytics streams the pre-tool text, the query_lap_times call, and the lap-by-lap answer; document comparison shows two search_documents calls in one round and the page-cited answer; booking streams the prefilled trip form
  • Invalid THESYS_API_KEY: /api/chat returns HTTP 401 with Gateway's message
  • Unknown OPENUI_MODEL (booking): /api/chat returns HTTP 502 with {"error":"Unsupported model: …"} instead of a 200 stream that fails; a normal request still streams (200, 709 chunks)
  • After simplifying the routes, in the browser: analytics answers the fastest-laps starter with a query_lap_times call, a table, and a chart, and after a reload the reopened thread loads its question, tool call, and answer; document comparison answers "Compare revenue and growth" with page-cited sources; booking returns the prefilled New York form. No server errors.
  • With history loaded from Gateway, in the browser: analytics answers a question and a follow-up, and the converted history for that thread is user, assistant with tool_calls, the matching tool result, and the answer, for each turn; document comparison extends a revenue and R&D comparison with headcount; booking goes from the prefilled form to stay cards to the summary. No server errors.
  • After simplifying: /api/chat returns 400 without a threadId or when the last message isn't the user's, and the token route returns 200 with Cache-Control: private, no-store
  • A token from the route creates conversations with user_id: demo-user and the cookbook's app_id
  • Probe: a token for another user_id or app_id gets 404 on the conversation and its items
  • After Gateway fixed the items endpoint: a reloaded booking thread loads its question and prefilled form from GET …/items (200)
  • Stopped runs, against real Gateway storage: cancelled before any answer stores the user's message; cancelled mid-answer (after 50 chunks) stores the user's message; stopped after a tool finished stores the message, tool call, and result; an unknown model stores the user's message
  • Docs site, locally at refactor(cookbooks): answer with the chat library #1264's head on the rebased stack: the three cookbook pages render with no console or server errors, each "Thread storage" callout is one sentence, the diagrams have no stage numbers, and no code block is taller than its 600px box
  • At this PR's head, no cookbook sample is over 30 lines, and Prettier passes on the three cookbook pages

🤖 Generated with Claude Code

@vercel

vercel Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
openui-docs Ready Ready Preview Sep 30, 2026 2:02pm UTC

Request Review

@vishxrad
vishxrad added this pull request to stack #1262 September 29, 2026 07:48
@vishxrad vishxrad changed the title feat(cookbooks): store threads with the Gateway Conversations API feat(cookbooks): store threads in Gateway conversations and run tools with runTools() Sep 29, 2026
@vishxrad
vishxrad force-pushed the visharad/cookbooks-gateway-conversations branch 2 times, most recently from 619b3a1 to 42ca1ce Compare September 29, 2026 11:28
@vishxrad
vishxrad force-pushed the visharad/cookbooks-gateway-conversations branch from 98a388a to e229607 Compare September 30, 2026 09:21
vishxrad and others added 22 commits September 30, 2026 17:53
Follow the self-hosted template: the chat route forwards the runTools()
runner's completion chunks as server-sent events, the format Gateway
streams, and Agent Interface reads them with openAIAdapter() instead of
AG-UI events built in the route. This drops the event mapping from each
route.

The route still waits for the first completion before responding, so a
rejected key or rate limit returns Gateway's HTTP status.

What changes for readers, and the docs now say so:
- Tool call arguments stream in again.
- Chat Completions has no chunk for a tool result, so Behind the scenes
  shows each call's arguments but not its result.
- A later error ends the stream, and Agent Interface reports that the
  request failed instead of showing Gateway's message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chat Completions doesn't write to a Gateway conversation, so each chat
route appends the turn itself once the model has answered: the user's
message, each tool call and its result, and the answer. It calls
storeChatCompletionHistory() from @openuidev/server with the threadId
that Agent Interface's cloud storage created, and passes only the new
turn, taken from runner.messages after the request's messages. A
stopped run or one without an answer isn't stored, and a failed write
is logged without affecting the answer.

The cookbook pages and READMEs describe the stored turns, the analytics
walkthrough shows the call, the Verify steps open a reloaded thread,
and the deployment notes say to check that the threadId belongs to the
signed-in user before storing.

Gateway currently answers the helper's POST /v1/conversations/{id}/items
with 404 "Conversation not found", so threads still reopen empty until
that is fixed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each chat route returns runTools()' own stream instead of re-encoding
every chunk as a server-sent event. toReadableStream() carries each
completion's chunks as one JSON object per line, so the pages read it
with openAIReadableStreamAdapter() instead of openAIAdapter(). The two
adapters handle chunks the same way and differ only in line framing.

Turn storage moves to runner.done(), and the route still waits for the
first completion before responding, so a rejected key or rate limit
keeps returning Gateway's HTTP status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e events

Build the turn from the user's message plus each message the runner
emits, instead of slicing runner.messages past the request's messages.
The message event is part of the runner's typed events and fires only
for messages the run adds, so the route no longer depends on
runner.messages starting with a copy of the request.

Drop the store-call walkthrough from the analytics page's Stream the
answer step. The storage paragraph in Connect Agent Interface already
names storeChatCompletionHistory(), as the other pages and READMEs do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chat Completions has no chunk for a tool result, so a live answer shows
each tool call without its result. The route now stores each turn with
its tool results, so say "in a live answer" on each cookbook page and
README instead of implying the results never appear.

The analytics Verify section said to check values against the tool
result, which readers can no longer see. Point to the race data in
data/f1.sqlite instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop the check that skipped storing a turn without a final answer.
Failed and stopped runs already reject runner.done() and are never
stored, and storeChatCompletionHistory() skips an assistant message
with no text. The only case the check caught was a run that hit
maxChatCompletions on a tool call. Storing that turn keeps the question
and its tool calls after a reload, as the person saw them live.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gateway reports some first-round errors, such as an unknown model,
inside a successful response. Waiting for the runner's connect event
let those through as a 200 stream that failed mid-way, and Agent
Interface showed "Failed to fetch". Waiting for the first chunk returns
them as an HTTP error with Gateway's message, falling back to 502 when
Gateway gives no status. A rejected key or rate limit still returns
Gateway's status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
toReadableStream() writes one JSON object per line, not server-sent
events, so return it as application/x-ndjson. Drop the Connection
header, which Node sets itself and HTTP/2 does not allow. The page's
openAIReadableStreamAdapter() reads the body without checking its type.

The 400 for an invalid request now names threadId, which the route
has required since it started storing turns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fails

Store the turn however the run ends. A stopped or failed run used to
skip storing, so the user's message disappeared when the thread was
opened again. The route now stores the user's message plus any tool
calls and results that finished; a partly streamed answer isn't kept,
because the runner adds an assistant message only when its completion
ends. A failed write is still logged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
parseChatRequest threw specific messages ("Send JSON.", "Request too
large.") that the route never showed, since it always answers an
invalid request with one 400. Return undefined instead, and check for
it with a plain if. body becomes a const, so the store callback reads
body.threadId directly. Wrong content type, invalid JSON, a body over
1 MB, and failed validation still return the same 400.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tions

- Say the model writes the answer in OpenUI Lang instead of calling it
  an OpenUI program.
- Add a How it works step and a Conversations box in the diagram: the
  server saves each turn to the thread's Gateway conversation, and the
  chat interface lists and reopens threads from there.
- Replace the stream, tool-loop, and error-handling details in Stream
  the answer with one paragraph that points to the chat route, and cut
  the storage paragraph down to what the reader needs.
- List the packages and APIs the example uses in a References section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Apply the analytics page's changes to the document comparison and
booking assistant pages:

- Add a How it works step for Gateway keeping the thread, and show the
  conversation in each diagram: a third "Keep the thread" stage in the
  comparison diagram, and a Conversation box fed by your server in the
  booking diagram.
- Cut the storage paragraph down to what the reader needs, and replace
  the tool-loop, stream, and request-limit details with a short
  paragraph that points to the chat route.
- Call the booking prompt's examples OpenUI Lang examples instead of
  programs.
- List the packages and APIs each example uses in a References section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Wrap the thread storage explanation on each cookbook page in a
Fumadocs info callout titled "Thread storage": how the browser reaches
Gateway conversations with a frontend token, and how the chat route
saves each turn. On the comparison and booking pages, the callout also
holds the useOpenuiCloudStorage() sample it introduces.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sample grew to 31 lines when it gained storage={storage}, which is
past Fumadocs' 600px code block height, so it scrolled inside its box
and clipped its last line. Split it into the props (theme, brand, and
starters) and the slots (sidebar, Documents page, and thread header),
each followed by the bullets that explain it. No cookbook page has a
code block over 30 lines now.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Say that the route runs the tool the model calls, and leave the OpenAI
SDK helper to the route's code sample and the References section.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Analytics: remove the paragraphs on rendering with componentLibrary
  and on the "final ten laps" follow-up. How it works, the step 3 and 4
  samples, Try this conversation, and Verify already cover both.
- Document comparison: drop "searches run in parallel" from the route
  paragraph; How it works and the prompt rules already say it.
- Booking: drop the booking link and form submission clauses from the
  route paragraph; step 4 and How it works already say them.
- Move the agent framework paragraph on the comparison and booking
  pages from Stream the answer and Connect Agent Interface to Adapt,
  where the analytics page has it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Code samples on the cookbook pages ran past the code box, which fits
about 86 characters at desktop width, so long lines scrolled sideways
or sat against the edge. Wrap them at 80 characters:

- Reformat the route, search, trivago, and Agent Interface samples at
  80 columns, and break the thread header onto separate lines with a
  blank line between the slot groups.
- Shorten two comments, split the date picker's description string,
  and drop the Welcome description from the analytics sample, where it
  is not what the step is about.
- Stop Prettier from reformatting code in the cookbook pages
  (embeddedLanguageFormatting: off), since it would widen them back to
  the repo's 100-column width. Prettier still formats the prose.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop the Conversations API, frontend token, and restStorage details from
the cookbook pages. Each page now says in one sentence how threads are
stored and links storeChatCompletionHistory(). The diagrams still show
the conversation, and the example READMEs keep the implementation notes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the guide links at the end of each page into its References list,
so readers find every package, API, and guide in one place.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The numbered steps under each diagram walk through a different sequence
than the diagram's stages, so number only the steps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The chat routes validated requests with a zod schema, read the body
through a 1 MB streaming reader, checked the page origin, and turned a
missing key or missing data into a 503. The other examples and the
openui-cloud template read the body with request.json() and trust the
local page, so do the same here.

Each chat route still forwards only user and assistant text, returns
400 without a threadId or a closing user message, waits for the first
chunk so Gateway errors return as HTTP errors, and stores the turn. The
token routes drop the origin check. The pages and READMEs no longer say
the routes accept requests only from their own page.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tion

The browser never receives tool results over Chat Completions, so the
history it sent back had tool calls without results, which Gateway
rejects. Stripping the calls kept follow-ups working but hid from the
model where its earlier answers came from.

The chat routes now use only the new question from the browser and
load the earlier turns, tool calls and results included, from the
thread's Gateway conversation. src/lib/gateway-history.ts converts the
stored items back to Chat Completions messages, the reverse of
storeChatCompletionHistory(), and drops tool calls without a result.
The response ends only after the turn is stored, so a follow-up finds
it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Comment thread docs/content/docs/cookbooks/document-comparison.mdx Outdated
@vishxrad
vishxrad merged commit 600ee4e into main Sep 30, 2026
5 checks passed
@vishxrad
vishxrad deleted the visharad/cookbooks-gateway-conversations branch September 30, 2026 14:03

This branch was successfully deployed

1 active deployment
Preview — 36992c1d Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants