feat(cookbooks): store threads in Gateway conversations and run tools with runTools() - #1258
Merged
Merged
Conversation
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
5 of 7 tasks
vishxrad
added this pull request to stack #1262
September 29, 2026 07:48
vishxrad
force-pushed
the
visharad/cookbooks-gateway-conversations
branch
2 times, most recently
from
September 29, 2026 11:28
619b3a1 to
42ca1ce
Compare
vishxrad
force-pushed
the
visharad/cookbooks-gateway-conversations
branch
from
September 30, 2026 09:21
98a388a to
e229607
Compare
Follow the self-hosted template: the chat route forwards the runTools() runner's completion chunks as server-sent events, the format Gateway streams, and Agent Interface reads them with openAIAdapter() instead of AG-UI events built in the route. This drops the event mapping from each route. The route still waits for the first completion before responding, so a rejected key or rate limit returns Gateway's HTTP status. What changes for readers, and the docs now say so: - Tool call arguments stream in again. - Chat Completions has no chunk for a tool result, so Behind the scenes shows each call's arguments but not its result. - A later error ends the stream, and Agent Interface reports that the request failed instead of showing Gateway's message. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chat Completions doesn't write to a Gateway conversation, so each chat
route appends the turn itself once the model has answered: the user's
message, each tool call and its result, and the answer. It calls
storeChatCompletionHistory() from @openuidev/server with the threadId
that Agent Interface's cloud storage created, and passes only the new
turn, taken from runner.messages after the request's messages. A
stopped run or one without an answer isn't stored, and a failed write
is logged without affecting the answer.
The cookbook pages and READMEs describe the stored turns, the analytics
walkthrough shows the call, the Verify steps open a reloaded thread,
and the deployment notes say to check that the threadId belongs to the
signed-in user before storing.
Gateway currently answers the helper's POST /v1/conversations/{id}/items
with 404 "Conversation not found", so threads still reopen empty until
that is fixed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each chat route returns runTools()' own stream instead of re-encoding every chunk as a server-sent event. toReadableStream() carries each completion's chunks as one JSON object per line, so the pages read it with openAIReadableStreamAdapter() instead of openAIAdapter(). The two adapters handle chunks the same way and differ only in line framing. Turn storage moves to runner.done(), and the route still waits for the first completion before responding, so a rejected key or rate limit keeps returning Gateway's HTTP status. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e events Build the turn from the user's message plus each message the runner emits, instead of slicing runner.messages past the request's messages. The message event is part of the runner's typed events and fires only for messages the run adds, so the route no longer depends on runner.messages starting with a copy of the request. Drop the store-call walkthrough from the analytics page's Stream the answer step. The storage paragraph in Connect Agent Interface already names storeChatCompletionHistory(), as the other pages and READMEs do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Chat Completions has no chunk for a tool result, so a live answer shows each tool call without its result. The route now stores each turn with its tool results, so say "in a live answer" on each cookbook page and README instead of implying the results never appear. The analytics Verify section said to check values against the tool result, which readers can no longer see. Point to the race data in data/f1.sqlite instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop the check that skipped storing a turn without a final answer. Failed and stopped runs already reject runner.done() and are never stored, and storeChatCompletionHistory() skips an assistant message with no text. The only case the check caught was a run that hit maxChatCompletions on a tool call. Storing that turn keeps the question and its tool calls after a reload, as the person saw them live. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Gateway reports some first-round errors, such as an unknown model, inside a successful response. Waiting for the runner's connect event let those through as a 200 stream that failed mid-way, and Agent Interface showed "Failed to fetch". Waiting for the first chunk returns them as an HTTP error with Gateway's message, falling back to 502 when Gateway gives no status. A rejected key or rate limit still returns Gateway's status. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
toReadableStream() writes one JSON object per line, not server-sent events, so return it as application/x-ndjson. Drop the Connection header, which Node sets itself and HTTP/2 does not allow. The page's openAIReadableStreamAdapter() reads the body without checking its type. The 400 for an invalid request now names threadId, which the route has required since it started storing turns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fails Store the turn however the run ends. A stopped or failed run used to skip storing, so the user's message disappeared when the thread was opened again. The route now stores the user's message plus any tool calls and results that finished; a partly streamed answer isn't kept, because the runner adds an assistant message only when its completion ends. A failed write is still logged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
parseChatRequest threw specific messages ("Send JSON.", "Request too
large.") that the route never showed, since it always answers an
invalid request with one 400. Return undefined instead, and check for
it with a plain if. body becomes a const, so the store callback reads
body.threadId directly. Wrong content type, invalid JSON, a body over
1 MB, and failed validation still return the same 400.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tions - Say the model writes the answer in OpenUI Lang instead of calling it an OpenUI program. - Add a How it works step and a Conversations box in the diagram: the server saves each turn to the thread's Gateway conversation, and the chat interface lists and reopens threads from there. - Replace the stream, tool-loop, and error-handling details in Stream the answer with one paragraph that points to the chat route, and cut the storage paragraph down to what the reader needs. - List the packages and APIs the example uses in a References section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Apply the analytics page's changes to the document comparison and booking assistant pages: - Add a How it works step for Gateway keeping the thread, and show the conversation in each diagram: a third "Keep the thread" stage in the comparison diagram, and a Conversation box fed by your server in the booking diagram. - Cut the storage paragraph down to what the reader needs, and replace the tool-loop, stream, and request-limit details with a short paragraph that points to the chat route. - Call the booking prompt's examples OpenUI Lang examples instead of programs. - List the packages and APIs each example uses in a References section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Wrap the thread storage explanation on each cookbook page in a Fumadocs info callout titled "Thread storage": how the browser reaches Gateway conversations with a frontend token, and how the chat route saves each turn. On the comparison and booking pages, the callout also holds the useOpenuiCloudStorage() sample it introduces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The sample grew to 31 lines when it gained storage={storage}, which is
past Fumadocs' 600px code block height, so it scrolled inside its box
and clipped its last line. Split it into the props (theme, brand, and
starters) and the slots (sidebar, Documents page, and thread header),
each followed by the bullets that explain it. No cookbook page has a
code block over 30 lines now.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Say that the route runs the tool the model calls, and leave the OpenAI SDK helper to the route's code sample and the References section. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Analytics: remove the paragraphs on rendering with componentLibrary and on the "final ten laps" follow-up. How it works, the step 3 and 4 samples, Try this conversation, and Verify already cover both. - Document comparison: drop "searches run in parallel" from the route paragraph; How it works and the prompt rules already say it. - Booking: drop the booking link and form submission clauses from the route paragraph; step 4 and How it works already say them. - Move the agent framework paragraph on the comparison and booking pages from Stream the answer and Connect Agent Interface to Adapt, where the analytics page has it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Code samples on the cookbook pages ran past the code box, which fits about 86 characters at desktop width, so long lines scrolled sideways or sat against the edge. Wrap them at 80 characters: - Reformat the route, search, trivago, and Agent Interface samples at 80 columns, and break the thread header onto separate lines with a blank line between the slot groups. - Shorten two comments, split the date picker's description string, and drop the Welcome description from the analytics sample, where it is not what the step is about. - Stop Prettier from reformatting code in the cookbook pages (embeddedLanguageFormatting: off), since it would widen them back to the repo's 100-column width. Prettier still formats the prose. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Drop the Conversations API, frontend token, and restStorage details from the cookbook pages. Each page now says in one sentence how threads are stored and links storeChatCompletionHistory(). The diagrams still show the conversation, and the example READMEs keep the implementation notes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the guide links at the end of each page into its References list, so readers find every package, API, and guide in one place. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The numbered steps under each diagram walk through a different sequence than the diagram's stages, so number only the steps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The chat routes validated requests with a zod schema, read the body through a 1 MB streaming reader, checked the page origin, and turned a missing key or missing data into a 503. The other examples and the openui-cloud template read the body with request.json() and trust the local page, so do the same here. Each chat route still forwards only user and assistant text, returns 400 without a threadId or a closing user message, waits for the first chunk so Gateway errors return as HTTP errors, and stores the turn. The token routes drop the origin check. The pages and READMEs no longer say the routes accept requests only from their own page. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tion The browser never receives tool results over Chat Completions, so the history it sent back had tool calls without results, which Gateway rejects. Stripping the calls kept follow-ups working but hid from the model where its earlier answers came from. The chat routes now use only the new question from the browser and load the earlier turns, tool calls and results included, from the thread's Gateway conversation. src/lib/gateway-history.ts converts the stored items back to Chat Completions messages, the reverse of storeChatCompletionHistory(), and drops tool calls without a result. The response ends only after the turn is stored, so a follow-up finds it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vishxrad
force-pushed
the
visharad/cookbooks-gateway-conversations
branch
from
September 30, 2026 12:25
d087980 to
2abec5a
Compare
abhithesys
reviewed
Sep 30, 2026
abhithesys
approved these changes
Sep 30, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR now includes #1261 (
runTools()and raw completion chunks), since both changed the same chat routes. #1264 (chat library) is stacked on top.Summary
All three cookbooks (conversational analytics, document comparison, and booking assistant) now keep their threads in Gateway conversations and run their tools with the OpenAI SDK's
runTools().Thread storage
storage={useOpenuiCloudStorage({ token: "/api/frontend-token", features: { artifact: false } })}. The browser lists, creates, renames, and deletes threads directly with the Conversations API./api/frontend-tokenroute mints a short-lived token withTHESYS_API_KEY. It usesDEMO_USER_ID(defaultdemo-user) andAPP_ID(default per cookbook, for examplebooking-assistant-cookbook), the same variables apps fromopenui createuse. The key stays on the server, and each cookbook sees only its own threads even when all three share a key.storeChatCompletionHistory()from@openuidev/server(new dependency, pinned to0.1.0). The turn is the user's message, each tool call and its result, and the answer, collected from the runner'smessageevents and stored oncerunner.done()settles. The conversation id is thethreadIdthat Agent Interface's storage created. A stopped or failed run still stores the user's message and any tool steps that finished, so the question stays after a reload; a partly streamed answer isn't kept. A failed write is logged without affecting the answer. The response ends only after the write, so a follow-up finds the turn.src/lib/gateway-history.tspages throughGET /v1/conversations/{id}/itemsand converts the items back to Chat Completions messages, the reverse ofstoreChatCompletionHistory(), dropping tool calls that have no result (a stopped run). The browser can't supply this history: it never receives tool results over Chat Completions, so its copy has tool calls without results, which Gateway rejects with a 400. This way the model sees where each earlier answer came from, andconversation()in the route is the place to add compaction later.Generation
src/lib/tool-loop.tsis gone. The chat route callsgateway.chat.completions.runTools()with the tool's JSON schema and its executor.runner.toReadableStream(), which carries every completion's chunks as one JSON object per line (application/x-ndjson), and the page reads it withopenAIReadableStreamAdapter()andopenAIMessageFormat..asResponse()isn't available here: it belongs to a singlecreate()call, whilerunTools()makes one request per round and reads each response itself to find the tool calls.openui-cloudtemplate, the routes read the body withrequest.json()and trust the local page. The chat route returns 400 without athreadIdor a closing user message. The zod request schema, 1 MB body reader, origin checks, and 503s for a missing key or missing data are gone; a missing key or unprepared data now throws.maxChatCompletionsis 3, 4, and 3. The route waits for the first chunk before responding, so a rejected key, rate limit, or unknown model returns an HTTP error with Gateway's message. Gateway reports an unknown model inside a successful response, so waiting only for the response let it through as a broken stream.Docs
storeChatCompletionHistory(), and each diagram shows the conversation. The pages don't describe the Conversations API or frontend tokens. The READMEs keep those details, including what to change before deploying: sign users in, mint tokens for the signed-in user, and check that athreadIdbelongs to that user before storing a turn in it.runTools()only in code and References, and drop paragraphs that repeated earlier steps. Each page ends with one References list of the packages, APIs, and guides the example uses. Mentions of authentication link to the Gateway authentication page, and the booking prompt no longer callssearch_staysa Responses tool..prettierrc.jsonturns off embedded-code formatting fordocs/content/docs/cookbooks/*.mdx, since Prettier would otherwise rewrap the samples at 100 characters. A fewlibrary.tsandsources.tsxsample lines here are still longer (up to 156 characters), because refactor(cookbooks): answer with the chat library #1264 replaces those samples.What changes for readers
runTools()ends the run when a tool throws, so each route returns the error to the model as the tool result.maxChatCompletionson a tool call ends without an answer.Gateway items endpoint (fixed)
storeChatCompletionHistory()posts toPOST /v1/conversations/{id}/items. Until 30 Sep 2026, Gateway answered valid requests with 404 "Conversation not found" for conversations that every other endpoint found. It now returns 200, the repro script prints "Not reproduced", and a reloaded thread loads its messages.Test plan
npm ciandnpm run verify(spec generation andnext buildwith type checking) in all three cookbooks at this PR's head, rebased on mainGET /v1/conversationsreturns 200, and the first message creates a conversation (201) and streams the answer into a new threadtoReadableStream(), in the browser: analytics streams the pre-tool text, thequery_lap_timescall, and the lap-by-lap answer; document comparison shows twosearch_documentscalls in one round and the page-cited answer; booking streams the prefilled trip formTHESYS_API_KEY:/api/chatreturns HTTP 401 with Gateway's messageOPENUI_MODEL(booking):/api/chatreturns HTTP 502 with{"error":"Unsupported model: …"}instead of a 200 stream that fails; a normal request still streams (200, 709 chunks)query_lap_timescall, a table, and a chart, and after a reload the reopened thread loads its question, tool call, and answer; document comparison answers "Compare revenue and growth" with page-cited sources; booking returns the prefilled New York form. No server errors.tool_calls, the matchingtoolresult, and the answer, for each turn; document comparison extends a revenue and R&D comparison with headcount; booking goes from the prefilled form to stay cards to the summary. No server errors./api/chatreturns 400 without athreadIdor when the last message isn't the user's, and the token route returns 200 withCache-Control: private, no-storeuser_id: demo-userand the cookbook'sapp_iduser_idorapp_idgets 404 on the conversation and its itemsGET …/items(200)🤖 Generated with Claude Code