diff --git a/SECURITY.md b/SECURITY.md index 3c9a024bda..0254ff0f4e 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -268,7 +268,7 @@ inputOrgId })`. It resolves from: route param slug → request **BYOK carve-out (intentional).** When a contributor sets a `BYOK_*` key ([editor/lib/ai/models.ts](editor/lib/ai/models.ts) — -`BYOK_OPENROUTER_API_KEY`, `BYOK_AI_GATEWAY_API_KEY`), `grida`/`model` +`BYOK_OPENROUTER_API_KEY`, `BYOK_VERCEL_AI_GATEWAY_API_KEY`), `grida`/`model` return a **bare** provider so the **AI-SDK text/chat path** bypasses the billing seam: no gate, no Metronome ingest, **and** the `MissingOrgIdError` runtime contract above does not fire (a bare @@ -284,8 +284,8 @@ actions still read the real balance and cannot silently drain credit while reporting `0`. **BYOK bypasses billing only — never auth.** `requireOrganizationId` and route/action auth always run, so a logged-in user with no resolvable org is still rejected. Gated solely by server-only, non-`NEXT_PUBLIC_` -env vars never set in the hosted product (same trust model as -`OPENAI_API_KEY` / `REPLICATE_API_TOKEN`). Fail-closed: `byok` is +env vars never set in the hosted product (same trust model as other +server-only credentials). Fail-closed: `byok` is `null` unless a key env var is a non-empty string, so any ambiguity falls back to the billed path. **Residual risk:** `byok` is resolved once at module load with no per-request guard — an accidental `BYOK_*` @@ -293,12 +293,26 @@ on a hosted/preview deploy would make every org bypass billing and the org-id sanity gate (auth still holds). Acceptable only because it is a contributor/self-host switch under the existing server-env trust model. +**Funded provider authority.** The shared Vercel AI Gateway provider reads only +`GG_VERCEL_AI_GATEWAY_API_KEY` for an explicit API key. An absent or blank value +retains the SDK's per-request platform OIDC resolution; an explicit empty SDK +setting prevents its ambient `AI_GATEWAY_API_KEY` lookup. Replicate predictions +require a nonblank `GG_REPLICATE_API_TOKEN` at execution, with no legacy token +fallback. Contributor overrides never supply these funded clients. Library query +embeddings still use the shared provider without user billing (or the contributor +override locally), and OpenAI model listing keeps its separate `OPENAI_API_KEY`. +Native CLI/Desktop credential names and stored provider IDs are independent of +this server configuration. + **Files bound by this id.** Run `grep -rn GRIDA-SEC-003 .` to enumerate. Today: - [editor/lib/auth/organization.ts](editor/lib/auth/organization.ts) — `requireOrganizationId`. - [editor/lib/ai/server.ts](editor/lib/ai/server.ts) — single seam entry; unconditional runtime gate; BYOK layer switch. -- [editor/lib/ai/models.ts](editor/lib/ai/models.ts) — BYOK layer (bare provider, bypasses billing). +- [editor/lib/ai/models.ts](editor/lib/ai/models.ts) — contributor BYOK and explicit funded Vercel AI Gateway/OIDC authority. +- [Provider credential tests](editor/lib/ai/__tests__/models.test.ts) and + [Replicate credential tests](editor/lib/ai/__tests__/replicate-credentials.test.ts) — + funding-role separation, per-request OIDC, missing credentials and preserved billing. - [GG Tripo execution](editor/lib/ai/gg-three-d.ts) and [billing contract tests](editor/lib/ai/__tests__/gg-three-d.test.ts) — verified org input, unconditional gate, infrastructure-key-only execution and actual receipts. @@ -2507,7 +2521,7 @@ generation receipts without a fixed projection. All selected file/environment/stdin keys use the shared AI provider admission policy: bounded to 4 KiB, normalized, header-safe, free of known template values, and checked against documented first-party formats without guessed suffix lengths. - Vercel legacy keys remain opaque where upstream specifies no retirement contract. + Vercel AI Gateway legacy keys remain opaque where upstream specifies no retirement contract. Stdin has a cancellable 30-second bound. Only the trusted SDK key reader receives secret strings. Status projects presence and source and plaintext storage mode; disposal drops private references. Explicit @@ -2566,7 +2580,7 @@ generation receipts without a fixed projection. 6. **Explicit credential checks before registration.** CLI `providers configure` validates the entered key through the shared AI owner, then invokes that owner's single authenticated GET before opening custody. The host permits only OpenRouter's - `/api/v1/key`, Vercel's `/v1/credits`, fal's `/v1/models/pricing` with exactly + `/api/v1/key`, Vercel AI Gateway's `/v1/credits`, fal's `/v1/models/pricing` with exactly one fixed `endpoint_id=fal-ai/flux/dev`, and Tripo's `/v3/account/balance`. This does not grant other platform APIs. The shared owner rejects redirects, bounds the request/body lifecycle to ten seconds diff --git a/docs/cli/providers.md b/docs/cli/providers.md index 6a73877b96..c81dbc1749 100644 --- a/docs/cli/providers.md +++ b/docs/cli/providers.md @@ -23,7 +23,7 @@ grida providers list ``` Enter your key at the hidden prompt. Every selected key passes static validation. -Configuration additionally checks OpenRouter, Vercel, and fal once before saving. +Configuration additionally checks OpenRouter, Vercel AI Gateway, and fal once before saving. A rejected or unavailable check leaves existing credentials unchanged. ElevenLabs has no suitable permission-neutral check and saves with verification marked `not_supported`. diff --git a/docs/contributing/billing.md b/docs/contributing/billing.md index aaa6fad10e..e34b28803b 100644 --- a/docs/contributing/billing.md +++ b/docs/contributing/billing.md @@ -1,4 +1,7 @@ --- +title: Contributing to Grida billing +description: Configure local billing sandboxes, contributor BYOK, and server provider credentials for Grida Gateway. +keywords: [grida, contributing, billing, credentials, byok, grida gateway] format: md --- @@ -21,18 +24,80 @@ If you are **not** working on the billing surface and only need the **AI chat / # editor/.env.local (gitignored) BYOK_OPENROUTER_API_KEY=sk-or-v1-... # https://openrouter.ai/keys # …or, if you have one, a dedicated Vercel AI Gateway key: -# BYOK_AI_GATEWAY_API_KEY=... +# BYOK_VERCEL_AI_GATEWAY_API_KEY=... ``` - Bypasses **billing only — never auth.** Still sign in (`insider@grida.co` / `password`); a resolvable org is still required (an unauthenticated request still 401s). -- **Text/chat only** — BYOK swaps the AI-SDK provider, so only the text path is unbilled. Image/audio go through Replicate (`withTransaction`) and **still gate + bill even under BYOK** — those features need the full billing setup. (OpenRouter also exposes no image/audio models.) Catalog model IDs are unchanged; use IDs your provider accepts (edit `editor/lib/ai/models.ts` locally if one 404s). -- Precedence if both are set: OpenRouter, then Vercel. Fail-closed — an empty/unset (or whitespace-only) key falls back to the billed path. +- **Text/chat only** — the contributor override swaps the text provider. Hosted image, video, audio, and 3D operations **still gate + bill even under contributor BYOK** and need the full billing setup. Catalog model IDs are unchanged; use IDs your provider accepts. +- Precedence if both are set: OpenRouter, then Vercel AI Gateway. An empty/unset (or whitespace-only) contributor key does not select BYOK; without another contributor key, text uses the billed path. - **Never set `BYOK_*` on a hosted or preview deploy.** It disables billing **and** the org-id sanity gate for every org. Contributor / self-host / local only. See [SECURITY.md](https://github.com/gridaco/grida/blob/main/SECURITY.md) (`GRIDA-SEC-003`, BYOK carve-out). Working on billing itself? Ignore BYOK and continue with the full setup below. --- +## Server provider credentials + +`GG_` identifies Grida-managed provider credentials used by the funded server +path. `BYOK_` identifies a contributor's server override. Neither prefix +identifies a deployment environment: scope secrets separately for Production, +Preview, and Development. + +| Consumer | Configuration | Selection | +| ----------------------------------------------- | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Funded Vercel AI Gateway text, image, and video | `GG_VERCEL_AI_GATEWAY_API_KEY` or platform `VERCEL_OIDC_TOKEN` | Use a nonblank GG key when configured; an unset or whitespace-only value selects platform OIDC. No fallback to `AI_GATEWAY_API_KEY` or contributor keys for funded authority. | +| Funded Replicate operations | `GG_REPLICATE_API_TOKEN` | Required and nonblank when a Replicate operation runs; no fallback to `REPLICATE_API_TOKEN` or BYOK. | +| Funded Tripo operations | `GG_TRIPO_API_KEY` | Unchanged; no generic or BYOK fallback. | +| Contributor text override | `BYOK_OPENROUTER_API_KEY`, then `BYOK_VERCEL_AI_GATEWAY_API_KEY` | Nonblank keys select the contributor's provider, bypassing billing only. | +| Library query embeddings | Shared contributor or Vercel AI Gateway provider | Remain unbilled internal operations with the existing contributor precedence. Sharing a funded provider credential does not add customer metering. | +| OpenAI model discovery | `OPENAI_API_KEY` | Unchanged; the model-list operation is nonbillable. | + +Vercel supplies `VERCEL_OIDC_TOKEN` through its platform authentication lifecycle. +Keep that variable and its lifecycle intact; an OIDC deployment does not need a +new static Vercel AI Gateway key for the naming convention. Never copy an OIDC +token into a static secret. Missing provider configuration must fail at the +affected operation without preventing unrelated routes from loading. + +These server names are separate from installed CLI/Desktop provider credentials. +The native provider ID remains `vercel`, and the CLI environment override remains +`AI_GATEWAY_API_KEY`. See [CLI provider keys](../cli/providers.md). The contributor +rename is a direct cutover: replace `BYOK_AI_GATEWAY_API_KEY` with +`BYOK_VERCEL_AI_GATEWAY_API_KEY` in local server configuration; the old name has +no compatibility alias. + +### Deploying the credential rename + +Code review and infrastructure changes are separate steps. The naming change +does not require rotating provider keys. + +1. **A — code and draft PR.** Review the reader/consumer mapping, environment + examples, build environment forwarding, and synthetic credential-selection + tests. Record the infrastructure prerequisite in the draft PR; local work + does not authorize deployment or secret changes. +2. **B1 — before merge, after explicit operator GO.** Inventory environment scopes, + shared variables, and branch overrides. Add `GG_REPLICATE_API_TOKEN` in each + scope that runs funded Replicate operations, retaining `REPLICATE_API_TOKEN` + for the running deployment and rollback. If a deployment already uses a + Grida-managed `AI_GATEWAY_API_KEY`, add its value as + `GG_VERCEL_AI_GATEWAY_API_KEY` before deploying the new reader. OIDC deployments + need no Vercel AI Gateway API-key addition. Verify the candidate's affected provider + execution, billing, and shared Library embeddings before merge. +3. **B2 — after merge, after explicit operator GO.** Verify the intended deployed + revision and its effective configuration. After the agreed rollback window, + check for remaining consumers and remove obsolete server variables only from + migrated scopes. Retain any old key still needed by a rollback deployment or + another consumer. Removing an environment entry does not revoke the provider + key; revocation needs its own consumer check. + +Treat secrets as values to transfer directly between approved stores, never as +review evidence. Record names and scopes without printing values. For write-only +secrets, an operator must enter the original value or issue a replacement while +retaining the old key through deployment verification. Validate the configuration +on a new deployment; editing project variables does not update an already +running deployment. + +--- + ## What you need - Local Supabase running (`supabase start`). @@ -215,4 +280,4 @@ User-facing billing copy: [`docs/platform/billing.mdx`](../platform/billing.mdx) | `WEBHOOK_TUNNEL_HOSTNAME` | `editor/.env.test.local` | | `BILLING_E2E`, `BILLING_TEST_MODE`, `APP_URL` | `editor/.env.test` (committed) | -**Contributor BYOK (alternative — not required):** `BYOK_OPENROUTER_API_KEY` or `BYOK_AI_GATEWAY_API_KEY` in `editor/.env.local`. When set, the AI seam bypasses billing entirely and **none** of the Metronome rows above are needed. Auth is still required. See [Just need AI to work?](#just-need-ai-to-work-byok-instead-no-billing-setup). +**Contributor BYOK (alternative — not required):** `BYOK_OPENROUTER_API_KEY` or `BYOK_VERCEL_AI_GATEWAY_API_KEY` in `editor/.env.local`. When set, text/chat bypasses billing and **none** of the Metronome rows above are needed for that path. Auth is still required. See [Just need AI to work?](#just-need-ai-to-work-byok-instead-no-billing-setup). diff --git a/docs/editor/desktop/chatgpt-subscription.md b/docs/editor/desktop/chatgpt-subscription.md index e623dd551f..3030afb1e3 100644 --- a/docs/editor/desktop/chatgpt-subscription.md +++ b/docs/editor/desktop/chatgpt-subscription.md @@ -57,7 +57,7 @@ The model picker organizes other text models by provider: - **Grida** is always shown. Its hosted models are metered against your organization's prepaid Grida AI credit. -- **OpenRouter** and **Vercel** appear after their keys are configured. +- **OpenRouter** and **Vercel AI Gateway** appear after their keys are configured. - **Ollama** appears after it is configured with at least one model. Other provider groups appear when they are configured and available. diff --git a/docs/editor/desktop/local-models.md b/docs/editor/desktop/local-models.md index a8650e1c3c..4e9749638e 100644 --- a/docs/editor/desktop/local-models.md +++ b/docs/editor/desktop/local-models.md @@ -20,7 +20,7 @@ machine, served by [Ollama](https://ollama.com). There is no account to create and no API key to paste — your prompts, files, and the model's responses never leave your computer. -You can use local models alongside provider keys (OpenRouter, Vercel), or +You can use local models alongside provider keys (OpenRouter, Vercel AI Gateway), or as your only setup. ## Requirements diff --git a/docs/models/index.md b/docs/models/index.md index 4e71fe62de..0d18f0fbf9 100644 --- a/docs/models/index.md +++ b/docs/models/index.md @@ -154,7 +154,7 @@ than applying GPT Image 2's per-image table. Its estimate excludes input charges; use actual provider usage when comparing total costs. fal rounds the total charge up to the nearest $0.0001. -[Vercel AI Gateway](https://vercel.com/ai-gateway/models/gpt-image-2.5-flare) publishes $5/M input tokens, $1.25/M cached input tokens, and $30/M output tokens. [OpenRouter](https://openrouter.ai/api/v1/images/models/openai/gpt-image-2.5-flare/endpoints) publishes $5/M text input, $8/M image input, and $30/M image output tokens. Grida keeps each provider's published meter separately. Hosted image billing uses Gateway's reported response cost when available. Without that receipt, it falls back to the catalog's coarse per-image estimate ($0.055 for these variants), which can differ from actual token cost across quality levels and dimensions. +[Vercel AI Gateway](https://vercel.com/ai-gateway/models/gpt-image-2.5-flare) publishes $5/M input tokens, $1.25/M cached input tokens, and $30/M output tokens. [OpenRouter](https://openrouter.ai/api/v1/images/models/openai/gpt-image-2.5-flare/endpoints) publishes $5/M text input, $8/M image input, and $30/M image output tokens. Grida keeps each provider's published meter separately. Hosted image billing uses Vercel AI Gateway's reported response cost when available. Without that receipt, it falls back to the catalog's coarse per-image estimate ($0.055 for these variants), which can differ from actual token cost across quality levels and dimensions. **GPT Image 2** (`openai/gpt-image-2`) — _deprecated in Grida, superseded by GPT Image 2.5_ @@ -187,7 +187,7 @@ Provider support is checked separately from the model's native capability. OpenAI added GPT Image 2 transparency in its [August 20 update](https://developers.openai.com/api/docs/changelog), and fal exposes it. OpenRouter's [GPT Image 2 endpoint](https://openrouter.ai/api/v1/images/models/openai/gpt-image-2/endpoints) and current GPT Image 2.5 endpoints still accept only `auto` or `opaque`. -Vercel transparency for GPT Image 2 remains unverified, although its +Vercel AI Gateway transparency for GPT Image 2 remains unverified, although its [Flare](https://vercel.com/ai-gateway/models/gpt-image-2.5-flare) and [Sunburst](https://vercel.com/ai-gateway/models/gpt-image-2.5-sunburst) pages explicitly support it. Grida only offers transparency on verified routes. @@ -305,8 +305,8 @@ which is cheaper on every provider; this is not an upstream Recraft retirement. ## Video Generation Models Video models are billed per second of generated output, by resolution and -whether audio is generated. The rates below are the Grida-hosted (Vercel -gateway) rates; the hosted route always generates the model's default audio +whether audio is generated. The rates below are the Grida-hosted (Vercel AI +Gateway) rates; the hosted route always generates the model's default audio mode, so the silent rates are informational until the request can carry an audio mode. @@ -323,7 +323,7 @@ provider meters one. Wan and Grok bundle audio into a single rate. `Seedance 2.0` and `Seedance 2.5` (`bytedance/seedance-2.0`, `-2.5`) are catalogued for bring-your-own-key use through fal, but are **not available on -the hosted route**: the gateway meters them per video token rather than per +the hosted route**: Vercel AI Gateway meters them per video token rather than per second, and there is no honest per-second conversion, so Grida cannot pre-price a hosted request. This will change when hosted video is metered after generation. `Seedance 2.5` is the newer generation but not a cheaper diff --git a/docs/wg/ai/agent/chatgpt-subscription-provider.md b/docs/wg/ai/agent/chatgpt-subscription-provider.md index 8204ee9541..deea6488c1 100644 --- a/docs/wg/ai/agent/chatgpt-subscription-provider.md +++ b/docs/wg/ai/agent/chatgpt-subscription-provider.md @@ -479,7 +479,7 @@ the highest-capability compatible model (`GPT-5.6 Sol`). This default is derived from live readiness and MUST NOT be persisted as a global preference. The provider group order is ChatGPT Subscription when ready, Grida, configured text BYOK providers, then configured compatible endpoints. The Grida, OpenRouter, -and Vercel groups may repeat catalog model ids because each tuple names a +and Vercel AI Gateway groups may repeat catalog model ids because each tuple names a different cost and privacy boundary. OpenAI-shaped model ids outside the closed ChatGPT set never gain ChatGPT eligibility from their name alone. diff --git a/docs/wg/cli/credential-custody.md b/docs/wg/cli/credential-custody.md index 8a40b1c3c2..1274fb2bda 100644 --- a/docs/wg/cli/credential-custody.md +++ b/docs/wg/cli/credential-custody.md @@ -166,7 +166,7 @@ claims about a provider's key length. The [shared provider policy](https://github.com/gridaco/grida/blob/main/packages/grida-ai/README.md) owns the exact rules, upstream references and authenticated check endpoints. -`providers configure` checks a newly entered OpenRouter, Vercel or fal key once +`providers configure` checks a newly entered OpenRouter, Vercel AI Gateway or fal key once before saving, including when input comes from `--key-stdin`. A rejected, permission-denied, timed-out or inconclusive check does not replace the old key. ElevenLabs has no suitable permission-neutral check; its key is saved with static diff --git a/docs/wg/platform/billing/known-issues.md b/docs/wg/platform/billing/known-issues.md index 00045cbacc..7c89d7ddc9 100644 --- a/docs/wg/platform/billing/known-issues.md +++ b/docs/wg/platform/billing/known-issues.md @@ -160,7 +160,7 @@ costs. A single average invocation estimate cannot represent every request; aggregate input/output token counts also cannot distinguish differently priced text and image tokens. -**Current behavior.** When Gateway supplies a valid response cost, hosted image +**Current behavior.** When Vercel AI Gateway supplies a valid response cost, hosted image usage is metered from that USD receipt, including a legitimate zero. Otherwise, the existing catalog estimate is used. For GPT Image 2.5 that fallback is $0.055 per requested image, regardless of quality and dimensions, and can diff --git a/docs/wg/platform/hosted-ai.md b/docs/wg/platform/hosted-ai.md index 3f6da30ee2..45124336d0 100644 --- a/docs/wg/platform/hosted-ai.md +++ b/docs/wg/platform/hosted-ai.md @@ -403,8 +403,8 @@ speed label, and the like are closed unions in TypeScript but are accepted as plain strings when parsed, and an unknown _provider_ binding is dropped rather than rejecting its card. Otherwise publishing a model from a new vendor, or adding a provider, would be a breaking publish requiring a -client release — the exact failure this system exists to remove. (Vercel's -AI Gateway learned this in public: it validated its model-kind field as a +client release — the exact failure this system exists to remove. (Vercel AI +Gateway learned this in public: it validated its model-kind field as a hard enum, so the day a new kind shipped the entire listing failed to parse for every client. It now accepts loosely and filters unknown rows.) diff --git a/editor/.env.example b/editor/.env.example index b1bacfceb1..65b4e975ca 100644 --- a/editor/.env.example +++ b/editor/.env.example @@ -41,19 +41,25 @@ NEXT_PUBLIC_GRIDA_LOCALHOST_REGION="us-west-1" # mapbox NEXT_PUBLIC_MAPBOX_ACCESS_TOKEN=... -# openai (used by provider-specific tools like webSearch; model selection is in lib/ai/models.ts) +# openai (nonbillable model listing; separate from GG-funded execution) OPENAI_API_KEY='sk-xxx' -# vercel ai gateway (optional; billed path — implicit AI_GATEWAY_API_KEY/OIDC) -# AI_GATEWAY_API_KEY=... +# GRIDA-GG: gateway — Grida-funded provider credentials (server-only). +# Vercel AI Gateway uses platform OIDC by default; keep VERCEL_OIDC_TOKEN as provided by Vercel. +# Optional explicit Grida-funded key; absent/blank preserves OIDC. No AI_GATEWAY_API_KEY fallback. +# GG_VERCEL_AI_GATEWAY_API_KEY="" +# Replicate-backed image tools and music require this key; no REPLICATE_API_TOKEN fallback. +# GG_REPLICATE_API_TOKEN="" # # byok (contributor-only; local testing). when ANY is set, AI calls # route through that provider and BYPASS the billing layer (no credit # gate, no metering). does NOT bypass auth (login + org still required). # text/chat only; catalog model IDs unchanged (use IDs the provider -# accepts). precedence: openrouter, then vercel. +# accepts). precedence: OpenRouter, then Vercel AI Gateway. # BYOK_OPENROUTER_API_KEY=sk-or-v1-xxx # https://openrouter.ai/keys -# BYOK_AI_GATEWAY_API_KEY=... # dedicated Vercel AI Gateway key +# BYOK_VERCEL_AI_GATEWAY_API_KEY=... # dedicated Vercel AI Gateway key +# Direct rename of the former BYOK_AI_GATEWAY_API_KEY; update local overrides. +# Installed CLI/Desktop credentials are separate; the CLI still accepts AI_GATEWAY_API_KEY. # hosted-AI scoped tokens (GRIDA-SEC-006) # HS256 secret for the desktop hosted-AI tokens minted by diff --git a/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.byok.test.ts b/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.byok.test.ts index e63703b718..2288d1cf8b 100644 --- a/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.byok.test.ts +++ b/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.byok.test.ts @@ -65,9 +65,11 @@ vi.mock("@/lib/ai/models", async (orig) => { // BYOK active: the seam's language provider is the BARE provider. byok: { languageModel: bareModel }, isByokActive: () => true, - gateway: { + vercelAiGateway: { languageModel: () => { - throw new Error("gateway must not be used under BYOK text path"); + throw new Error( + "Vercel AI Gateway must not be used under BYOK text path" + ); }, imageModel: () => { throw new Error("unused"); diff --git a/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.test.ts b/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.test.ts index 099c513383..57b5450a10 100644 --- a/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.test.ts +++ b/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.test.ts @@ -57,13 +57,13 @@ vi.mock("@/lib/auth/organization", () => ({ requireOrganizationId: vi.fn<(...args: never[]) => unknown>(), })); -// Fake gateway: scripted V3 model behind the REAL billing middleware -// (server.ts wraps whatever `gateway` this module exports). +// Fake Vercel AI Gateway: scripted V3 model behind the REAL billing middleware +// (server.ts wraps whatever `vercelAiGateway` this module exports). vi.mock("@/lib/ai/models", async (orig) => { const real = await orig(); const fakeLanguageModel = (modelId: string) => ({ specificationVersion: "v3" as const, - provider: "fake-gateway", + provider: "fake-vercel-ai-gateway", modelId, supportedUrls: {}, doGenerate: async (options: unknown) => { @@ -87,7 +87,7 @@ vi.mock("@/lib/ai/models", async (orig) => { ...real, byok: null, isByokActive: () => false, - gateway: { + vercelAiGateway: { languageModel: fakeLanguageModel, imageModel: () => { throw new Error("unused in this suite"); diff --git a/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.ts b/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.ts index 5a9715d4e9..dc00df6fff 100644 --- a/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.ts +++ b/editor/app/(api)/(public)/api/v1/ai/chat/completions/route.ts @@ -56,7 +56,7 @@ export async function POST(request: Request) { if (!p.ok) return p.res; const req = p.data; - // Allowlist BEFORE any provider call — the gateway accepts arbitrary + // Allowlist BEFORE any provider call — Vercel AI Gateway accepts arbitrary // ids; an unlisted one must 404 here, not 500 at cost-card lookup. if (!isHostedTextModel(req.model)) return modelNotFound(req.model); diff --git a/editor/app/(api)/(public)/api/v1/ai/images/generations/route.test.ts b/editor/app/(api)/(public)/api/v1/ai/images/generations/route.test.ts index 1d0785a4b2..15bca0cc17 100644 --- a/editor/app/(api)/(public)/api/v1/ai/images/generations/route.test.ts +++ b/editor/app/(api)/(public)/api/v1/ai/images/generations/route.test.ts @@ -2,7 +2,7 @@ // GRIDA-GG: gateway — see docs/wg/platform/hosted-ai.md /** * POST /api/v1/ai/images/generations — token-gated hosted image - * generation through the REAL image billing middleware (fake gateway + * generation through the REAL image billing middleware (fake Vercel AI Gateway * model): pre-priced mills reach ingest, blocked orgs 402 before the * provider, unknown/unbound models 404, base64 results in the shared * protocol shape, and NO library upload (the daemon owns persistence). @@ -70,13 +70,13 @@ vi.mock("@/lib/ai/models", async (orig) => { ...real, byok: null, isByokActive: () => false, - gateway: { + vercelAiGateway: { languageModel: () => { throw new Error("unused"); }, imageModel: (modelId: string) => ({ specificationVersion: "v3" as const, - provider: "fake-gateway", + provider: "fake-vercel-ai-gateway", modelId, maxImagesPerCall: 4, doGenerate: async (options: unknown) => { @@ -112,7 +112,7 @@ const mockedIngest = vi.mocked(ingestUsageEvent); const SECRET = "images-secret-0123456789abcdef0123456789abcdef"; -// A real listed card with a vercel binding — resolved dynamically so +// A real listed card with a Vercel AI Gateway binding — resolved dynamically so // catalog churn doesn't rot the test. const CARD = ai.image .listed_models() @@ -384,7 +384,7 @@ describe("POST /api/v1/ai/images/generations", () => { expect(mockedIngest).not.toHaveBeenCalled(); }); - it("refuses a model without a Vercel binding before billing", async () => { + it("refuses a model without a Vercel AI Gateway binding before billing", async () => { vi.mocked(ai.image.binding).mockReturnValueOnce(null); const { token } = await signGgToken("user-1", 7); const res = await POST(request({ model_id: CARD.id, prompt: "x" }, token)); diff --git a/editor/app/(api)/(public)/api/v1/ai/images/generations/route.ts b/editor/app/(api)/(public)/api/v1/ai/images/generations/route.ts index 59f2b432d0..4628c2fba8 100644 --- a/editor/app/(api)/(public)/api/v1/ai/images/generations/route.ts +++ b/editor/app/(api)/(public)/api/v1/ai/images/generations/route.ts @@ -10,8 +10,8 @@ * protocol carries no reference images; the sidecar resolver routes * i2i to BYOK providers). * - * Auth: scoped AI token only (GRIDA-SEC-006). Billing: Gateway response - * receipt when available, otherwise the shared `computeImageCostMills` + * Auth: scoped AI token only (GRIDA-SEC-006). Billing: Vercel AI Gateway + * response receipt when available, otherwise the shared `computeImageCostMills` * fallback (`providerOptions.grida.costMills`). NO library * upload — the daemon owns persistence. */ @@ -66,8 +66,8 @@ export async function POST(request: Request) { const card = ai.image.findImageModelCard(req.model_id); if (!card) return modelNotFound(req.model_id); - // Hosted serving goes through the gateway — a card without a - // vercel binding is not servable here (the sidecar's BYOK + // Hosted serving goes through Vercel AI Gateway — a card without a + // matching binding is not servable here (the sidecar's BYOK // adapters cover the rest). const binding = ai.image.binding(card, "vercel"); if (!binding) return modelNotFound(req.model_id); @@ -80,7 +80,8 @@ export async function POST(request: Request) { } // An explicit mode must never become an opaque, billed success when the - // gateway has not exposed the native control. Auto is the legacy default. + // Vercel AI Gateway binding has not exposed the native control. + // Auto is the legacy default. const background = req.background === "auto" ? undefined : req.background; if (background && !ai.image.supportsTransparentBackground(card, "vercel")) { return invalidRequest( diff --git a/editor/app/(api)/(public)/api/v1/ai/models/route.test.ts b/editor/app/(api)/(public)/api/v1/ai/models/route.test.ts index f3f2ab0b0d..303cb97a9b 100644 --- a/editor/app/(api)/(public)/api/v1/ai/models/route.test.ts +++ b/editor/app/(api)/(public)/api/v1/ai/models/route.test.ts @@ -3,7 +3,7 @@ /** * GET /api/v1/ai/models — token-gated allowlist composed from the ONE * catalog: all text entries (deprecated flagged), image/video cards - * with vercel bindings and listed Tripo generation/rigging models; + * with Vercel AI Gateway bindings and listed Tripo generation/rigging models; * tier annotation from the reverse tier map; no * pricing fields anywhere in the payload. */ @@ -46,7 +46,7 @@ describe("GET /api/v1/ai/models", () => { expect((await GET(request("junk"))).status).toBe(401); }); - it("includes both verified GPT Image 2.5 Gateway bindings", async () => { + it("includes both verified GPT Image 2.5 Vercel AI Gateway bindings", async () => { const { token } = await signGgToken("user-1", 7); const res = await GET(request(token)); expect(res.status).toBe(200); diff --git a/editor/app/(api)/private/ai/chat/route.ts b/editor/app/(api)/private/ai/chat/route.ts index 8ea71f5f61..0829dad858 100644 --- a/editor/app/(api)/private/ai/chat/route.ts +++ b/editor/app/(api)/private/ai/chat/route.ts @@ -74,7 +74,7 @@ export async function POST(req: NextRequest) { return { totalUsage: part.totalUsage, lastStepUsage, - // Send the gateway model ID (e.g. "openai/gpt-5-mini") so + // Send the catalog model ID (e.g. "openai/gpt-5-mini") so // tokenlens in can resolve cost. Fall back to the // raw provider ID if the spec isn't found. modelId: spec?.id ?? lastModelId, diff --git a/editor/app/(www)/(ai)/ai/_page.tsx b/editor/app/(www)/(ai)/ai/_page.tsx index c266415cc7..4912abb732 100644 --- a/editor/app/(www)/(ai)/ai/_page.tsx +++ b/editor/app/(www)/(ai)/ai/_page.tsx @@ -75,7 +75,8 @@ const markdown = { // for specific call sites. // // `@grida/ai-models/grida` is a pure data entry; `@/lib/ai/models` is NOT -// imported here because it carries the server-only gateway/BYOK seam. +// imported here because it carries the server-only Vercel AI Gateway/BYOK +// provider seam. // --------------------------------------------------------------------------- type ModelOption = { id: string; diff --git a/editor/app/(www)/(ai)/ai/models/page.tsx b/editor/app/(www)/(ai)/ai/models/page.tsx index d047f4340a..702c56387b 100644 --- a/editor/app/(www)/(ai)/ai/models/page.tsx +++ b/editor/app/(www)/(ai)/ai/models/page.tsx @@ -859,7 +859,7 @@ function VideoModelOutput({ model }: { model: AITypes.video.VideoModelCard }) { } const VideoProviderLabels: Record = { - vercel: "Vercel", + vercel: "Vercel AI Gateway", fal: "fal", openrouter: "OpenRouter", }; diff --git a/editor/app/desktop/settings/_components/media-model-readiness.test.ts b/editor/app/desktop/settings/_components/media-model-readiness.test.ts index 77406aa498..2d11d3d8a7 100644 --- a/editor/app/desktop/settings/_components/media-model-readiness.test.ts +++ b/editor/app/desktop/settings/_components/media-model-readiness.test.ts @@ -32,7 +32,7 @@ describe("MediaModelReadiness.visual", () => { ).toBe(false); }); - it("admits hosted media only for a Vercel-backed model", () => { + it("admits hosted media only for a model served by Vercel AI Gateway", () => { expect(MediaModelReadiness.visual({ providers }, new Set(), true)).toBe( true ); diff --git a/editor/app/desktop/settings/_components/media-model-readiness.ts b/editor/app/desktop/settings/_components/media-model-readiness.ts index 21517df92e..8cea6b4205 100644 --- a/editor/app/desktop/settings/_components/media-model-readiness.ts +++ b/editor/app/desktop/settings/_components/media-model-readiness.ts @@ -15,8 +15,8 @@ export namespace MediaModelReadiness { return false; } /** - * Hosted image/video resolution follows the catalogue's Vercel binding, - * while BYOK resolution intersects every connected provider with the exact + * Hosted image/video resolution follows the catalogue's Vercel AI Gateway + * binding, while BYOK resolution intersects every connected provider with the exact * bindings on this card. A pending source keeps the result pending unless * the other source has already proved the model runnable. */ diff --git a/editor/grida-canvas-hosted/ai/types.ts b/editor/grida-canvas-hosted/ai/types.ts index 4b7de1d4c4..c40eae950f 100644 --- a/editor/grida-canvas-hosted/ai/types.ts +++ b/editor/grida-canvas-hosted/ai/types.ts @@ -19,7 +19,7 @@ export type AgentMessageMetadata = { * model saw, which is the correct proxy for context window consumption. */ lastStepUsage?: LanguageModelUsage; - /** The model ID that produced this response (gateway format). */ + /** The catalog model ID, or the upstream response ID when unrecognized. */ modelId?: string; /** Maximum context window in tokens for this model. */ contextWindow?: number; diff --git a/editor/lib/ai/__tests__/generate-video.test.ts b/editor/lib/ai/__tests__/generate-video.test.ts index 4d909315c4..c8604af24c 100644 --- a/editor/lib/ai/__tests__/generate-video.test.ts +++ b/editor/lib/ai/__tests__/generate-video.test.ts @@ -76,7 +76,7 @@ beforeEach(() => { }); describe("methods.generateVideo", () => { - it("happy path: gates, generates via the vercel binding, ingests rate×duration", async () => { + it("happy path: gates, generates via the Vercel AI Gateway binding, ingests rate×duration", async () => { const duration = CARD.default.duration; const result = await methods.generateVideo(ORG, { model_id: MODEL_ID, diff --git a/editor/lib/ai/__tests__/models.test.ts b/editor/lib/ai/__tests__/models.test.ts new file mode 100644 index 0000000000..6623e0bece --- /dev/null +++ b/editor/lib/ai/__tests__/models.test.ts @@ -0,0 +1,205 @@ +// GRIDA-SEC-003 — contributor overrides and funded provider authority. +// GRIDA-GG: gateway — actual SDK authentication with synthetic requests only. +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; + +const billing = vi.hoisted(() => ({ + getEntitlement: + vi.fn(), + ingestUsageEvent: + vi.fn(), +})); + +vi.mock("@/lib/billing/metronome", () => ({ + ...billing, + refreshBalance: + vi.fn(), + BillingMetronomeError: class extends Error {}, +})); +vi.mock("@/lib/supabase/server", () => ({ + createLibraryClient: + vi.fn(), +})); +vi.mock("@/lib/auth/organization", () => ({ + requireOrganizationId: + vi.fn(), +})); + +const requestContext = Symbol.for("@vercel/request-context"); +const requests: { url: string; headers: Headers }[] = []; + +// The SDK checks token expiry locally; this unsigned fixture never leaves fetch. +function oidcToken(subject: string): string { + const encode = (value: unknown) => + Buffer.from(JSON.stringify(value)).toString("base64url"); + return `${encode({ alg: "none" })}.${encode({ + sub: subject, + exp: Math.floor(Date.now() / 1000) + 3600, + })}.fixture`; +} + +beforeEach(() => { + vi.resetModules(); + vi.clearAllMocks(); + requests.length = 0; + for (const key of [ + "GG_VERCEL_AI_GATEWAY_API_KEY", + "AI_GATEWAY_API_KEY", + "VERCEL_OIDC_TOKEN", + "BYOK_OPENROUTER_API_KEY", + "BYOK_VERCEL_AI_GATEWAY_API_KEY", + "BYOK_AI_GATEWAY_API_KEY", + ]) { + vi.stubEnv(key, undefined); + } + vi.stubGlobal(requestContext, undefined); + vi.stubGlobal( + "fetch", + vi.fn(async (input: string | URL | Request, init?: RequestInit) => { + const url = input instanceof Request ? input.url : String(input); + requests.push({ url, headers: new Headers(init?.headers) }); + if (!url.startsWith("https://ai-gateway.vercel.sh/")) { + throw new Error("Unexpected request in synthetic provider test"); + } + return Response.json({ + content: [{ type: "text", text: "fixture" }], + finishReason: { unified: "stop", raw: "stop" }, + usage: { + tokens: 1, + inputTokens: { total: 1 }, + outputTokens: { total: 1 }, + }, + embeddings: [[0.25, 0.5]], + }); + }) + ); +}); + +afterEach(() => { + vi.unstubAllGlobals(); + vi.unstubAllEnvs(); +}); + +describe("funded Vercel AI Gateway authority", () => { + it("uses only the explicit GG key even when other funding credentials are present", async () => { + vi.stubEnv("GG_VERCEL_AI_GATEWAY_API_KEY", " fixture-gg-key "); + vi.stubEnv("AI_GATEWAY_API_KEY", "fixture-native-key"); + vi.stubEnv("BYOK_VERCEL_AI_GATEWAY_API_KEY", "fixture-contributor-key"); + vi.stubEnv("VERCEL_OIDC_TOKEN", oidcToken("platform")); + const { vercelAiGateway } = await import("../models"); + + await vercelAiGateway("fixture/model").doGenerate({ prompt: [] }); + + expect(requests).toHaveLength(1); + expect(requests[0].headers.get("authorization")).toBe( + "Bearer fixture-gg-key" + ); + expect(requests[0].headers.get("ai-gateway-auth-method")).toBe("api-key"); + expect(requests[0].headers.get("http-referer")).toBe("https://grida.co"); + expect(requests[0].headers.get("x-title")).toBe("Grida"); + }); + + it.each([undefined, " \t "])( + "retains per-request platform OIDC with GG key %j and ignores an ambient API key", + async (key) => { + vi.stubEnv("GG_VERCEL_AI_GATEWAY_API_KEY", key); + vi.stubEnv("AI_GATEWAY_API_KEY", "fixture-unrelated-key"); + const firstToken = oidcToken("first-request"); + const secondToken = oidcToken("second-request"); + let token = firstToken; + vi.stubGlobal(requestContext, { + get: () => ({ headers: { "x-vercel-oidc-token": token } }), + }); + const { vercelAiGateway } = await import("../models"); + const model = vercelAiGateway("fixture/model"); + + await model.doGenerate({ prompt: [] }); + token = secondToken; + await model.doGenerate({ prompt: [] }); + + expect( + requests.map(({ headers }) => headers.get("authorization")) + ).toEqual([`Bearer ${firstToken}`, `Bearer ${secondToken}`]); + expect( + requests.every( + ({ headers }) => headers.get("ai-gateway-auth-method") === "oidc" + ) + ).toBe(true); + } + ); + + it("supports the platform OIDC environment credential", async () => { + const token = oidcToken("environment"); + vi.stubEnv("VERCEL_OIDC_TOKEN", token); + const { vercelAiGateway } = await import("../models"); + await vercelAiGateway("fixture/model").doGenerate({ prompt: [] }); + expect(requests[0].headers.get("authorization")).toBe(`Bearer ${token}`); + expect(requests[0].headers.get("ai-gateway-auth-method")).toBe("oidc"); + }); + + it("rejects unusable OIDC instead of dispatching with a generic key", async () => { + vi.stubEnv("AI_GATEWAY_API_KEY", "fixture-unrelated-key"); + // Malformed token parsing fails before the SDK's local credential refresh. + vi.stubEnv("VERCEL_OIDC_TOKEN", "fixture-invalid-oidc"); + const { vercelAiGateway } = await import("../models"); + await expect( + vercelAiGateway("fixture/model").doGenerate({ prompt: [] }) + ).rejects.toThrow(/authentication failed/i); + expect(requests).toHaveLength(0); + }); + + it("keeps Library embeddings on the shared provider without user billing", async () => { + const token = oidcToken("library"); + vi.stubEnv("VERCEL_OIDC_TOKEN", token); + const { embedTextUnbilled } = await import("../server"); + + await expect( + embedTextUnbilled("fixture/embedding", "query") + ).resolves.toEqual([0.25, 0.5]); + + expect(requests[0].headers.get("authorization")).toBe(`Bearer ${token}`); + expect(billing.getEntitlement).not.toHaveBeenCalled(); + expect(billing.ingestUsageEvent).not.toHaveBeenCalled(); + }); +}); + +describe("contributor Vercel AI Gateway override", () => { + it("uses the renamed, trimmed key independently of funded authority", async () => { + vi.stubEnv("BYOK_VERCEL_AI_GATEWAY_API_KEY", " fixture-contributor-key "); + const { byok, isByokActive } = await import("../models"); + expect(isByokActive()).toBe(true); + await byok!.languageModel("fixture/model").doGenerate({ prompt: [] }); + expect(requests[0].headers.get("authorization")).toBe( + "Bearer fixture-contributor-key" + ); + }); + + it("preserves OpenRouter-first selection", async () => { + vi.stubEnv("BYOK_OPENROUTER_API_KEY", " fixture-openrouter-key "); + vi.stubEnv("BYOK_VERCEL_AI_GATEWAY_API_KEY", "fixture-contributor-key"); + const { byok } = await import("../models"); + expect(byok!.languageModel("fixture/model").provider).toBe( + "openrouter.chat" + ); + }); + + it("ignores a blank OpenRouter override before selecting Vercel AI Gateway", async () => { + vi.stubEnv("BYOK_OPENROUTER_API_KEY", " \t "); + vi.stubEnv("BYOK_VERCEL_AI_GATEWAY_API_KEY", "fixture-contributor-key"); + const { byok } = await import("../models"); + await byok!.languageModel("fixture/model").doGenerate({ prompt: [] }); + expect(requests[0].headers.get("authorization")).toBe( + "Bearer fixture-contributor-key" + ); + }); + + it.each([undefined, " \t "])( + "does not activate a missing or blank new override (%j)", + async (key) => { + vi.stubEnv("BYOK_VERCEL_AI_GATEWAY_API_KEY", key); + vi.stubEnv("BYOK_AI_GATEWAY_API_KEY", "fixture-retired-name"); + const { byok, isByokActive } = await import("../models"); + expect(byok).toBeNull(); + expect(isByokActive()).toBe(false); + } + ); +}); diff --git a/editor/lib/ai/__tests__/replicate-credentials.test.ts b/editor/lib/ai/__tests__/replicate-credentials.test.ts new file mode 100644 index 0000000000..e67b59895e --- /dev/null +++ b/editor/lib/ai/__tests__/replicate-credentials.test.ts @@ -0,0 +1,229 @@ +// GRIDA-SEC-003 — funded Replicate credentials never borrow contributor authority. +// GRIDA-GG: gateway — provider selection through the real billing seam. +// GRIDA-EE: billing — entitlement and usage calls stay in the shared seam. +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; + +const h = vi.hoisted(() => ({ + events: [] as string[], + replicateOptions: vi.fn<(options: unknown) => void>(), + run: vi.fn< + ( + model: string, + options: { input: Record } + ) => Promise + >(), + getEntitlement: vi.fn< + (organizationId: number) => Promise<{ + allowed: boolean; + cachedBalanceCents: number; + reason?: string; + }> + >(), + ingestUsageEvent: + vi.fn< + ( + organizationId: number, + costMills: number, + options: { transactionId: string } + ) => Promise + >(), + openAiOptions: vi.fn<(options: unknown) => void>(), + listOpenAiModels: vi.fn<() => Promise<{ data: { id: string }[] }>>(), +})); + +vi.mock("replicate", () => ({ + default: class ReplicateMock { + readonly run = h.run; + constructor(options: unknown) { + h.replicateOptions(options); + } + }, +})); + +vi.mock("openai", () => ({ + default: class OpenAiMock { + readonly models = { list: h.listOpenAiModels }; + constructor(options: unknown) { + h.openAiOptions(options); + } + }, +})); + +vi.mock("@/lib/billing/metronome", () => ({ + getEntitlement: h.getEntitlement, + ingestUsageEvent: h.ingestUsageEvent, + refreshBalance: + vi.fn(), + BillingMetronomeError: class BillingMetronomeError extends Error { + constructor( + message: string, + readonly code: string, + readonly status = 500 + ) { + super(message); + } + }, +})); + +vi.mock("@/lib/supabase/server", () => ({ + createLibraryClient: + vi.fn(), +})); +vi.mock("@/lib/auth/organization", () => ({ + requireOrganizationId: + vi.fn(), +})); + +beforeEach(() => { + // Reset the actual server module so its lazy provider cache cannot carry + // a previous test's admitted credential into a missing-key case. + vi.resetModules(); + vi.resetAllMocks(); + h.events.length = 0; + for (const key of [ + "GG_REPLICATE_API_TOKEN", + "REPLICATE_API_TOKEN", + "BYOK_REPLICATE_API_TOKEN", + "GG_VERCEL_AI_GATEWAY_API_KEY", + "AI_GATEWAY_API_KEY", + "VERCEL_OIDC_TOKEN", + "BYOK_OPENROUTER_API_KEY", + "BYOK_AI_GATEWAY_API_KEY", + "BYOK_VERCEL_AI_GATEWAY_API_KEY", + "OPENAI_API_KEY", + ]) { + vi.stubEnv(key, undefined); + } + vi.stubGlobal( + "fetch", + vi.fn(() => Promise.reject(new Error("Unexpected network request"))) + ); + h.getEntitlement.mockImplementation(async () => { + h.events.push("gate"); + return { allowed: true, cachedBalanceCents: 1000 }; + }); + h.run.mockImplementation(async () => { + h.events.push("provider"); + return "https://example.test/music.mp3"; + }); + h.ingestUsageEvent.mockImplementation(async () => { + h.events.push("ingest"); + }); +}); + +afterEach(() => { + vi.unstubAllEnvs(); + vi.unstubAllGlobals(); +}); + +describe("funded Replicate credential admission", () => { + it("uses the trimmed GG token and still gates and meters under contributor BYOK", async () => { + vi.stubEnv("GG_REPLICATE_API_TOKEN", " synthetic-gg-replicate \n"); + vi.stubEnv("REPLICATE_API_TOKEN", "synthetic-legacy-replicate"); + vi.stubEnv("BYOK_REPLICATE_API_TOKEN", "synthetic-byok-replicate"); + vi.stubEnv("BYOK_VERCEL_AI_GATEWAY_API_KEY", "synthetic-contributor-text"); + const { methods, isByokActive } = await import("../server"); + + expect(isByokActive()).toBe(true); + await expect( + methods.generateMusic(7, "google/lyria-3", { prompt: "Quiet piano" }) + ).resolves.toEqual({ url: "https://example.test/music.mp3" }); + + expect(h.replicateOptions).toHaveBeenCalledExactlyOnceWith({ + auth: "synthetic-gg-replicate", + useFileOutput: false, + }); + expect(h.run).toHaveBeenCalledExactlyOnceWith("google/lyria-3", { + input: { prompt: "Quiet piano" }, + }); + expect(h.getEntitlement).toHaveBeenCalledExactlyOnceWith(7); + // Lyria 3 is sold at the catalog's $0.04 per generation: 40 mills. + expect(h.ingestUsageEvent).toHaveBeenCalledExactlyOnceWith(7, 40, { + transactionId: expect.any(String), + }); + expect(h.events).toEqual(["gate", "provider", "ingest"]); + }); + + it.each([ + { + name: "missing configuration", + gg: undefined, + legacy: undefined, + byok: undefined, + }, + { + name: "only the legacy token", + gg: undefined, + legacy: "synthetic-legacy", + byok: undefined, + }, + { + name: "only a BYOK token", + gg: undefined, + legacy: undefined, + byok: "synthetic-byok", + }, + { + name: "a blank GG token with conflicting tokens", + gg: " \t\n", + legacy: "synthetic-legacy", + byok: "synthetic-byok", + }, + ])( + "refuses $name before provider dispatch or ingestion", + async ({ gg, legacy, byok }) => { + vi.stubEnv("GG_REPLICATE_API_TOKEN", gg); + vi.stubEnv("REPLICATE_API_TOKEN", legacy); + vi.stubEnv("BYOK_REPLICATE_API_TOKEN", byok); + const { methods } = await import("../server"); + + await expect( + methods.generateMusic(7, "google/lyria-3", { prompt: "Quiet piano" }) + ).rejects.toThrow("GG_REPLICATE_API_TOKEN is not set"); + + expect(h.replicateOptions).not.toHaveBeenCalled(); + expect(h.run).not.toHaveBeenCalled(); + expect(h.ingestUsageEvent).not.toHaveBeenCalled(); + expect(h.events).toEqual(["gate"]); + } + ); + + it("does not construct or dispatch the provider when credit admission fails", async () => { + vi.stubEnv("GG_REPLICATE_API_TOKEN", "synthetic-gg-replicate"); + h.getEntitlement.mockResolvedValueOnce({ + allowed: false, + cachedBalanceCents: 0, + reason: "insufficient_credits", + }); + const { methods } = await import("../server"); + + await expect( + methods.generateMusic(7, "google/lyria-3", { prompt: "Quiet piano" }) + ).rejects.toMatchObject({ code: "blocked", status: 402 }); + + expect(h.getEntitlement).toHaveBeenCalledExactlyOnceWith(7); + expect(h.replicateOptions).not.toHaveBeenCalled(); + expect(h.run).not.toHaveBeenCalled(); + expect(h.ingestUsageEvent).not.toHaveBeenCalled(); + }); + + it("does not require Replicate configuration to list unbilled OpenAI models", async () => { + vi.stubEnv("OPENAI_API_KEY", "synthetic-openai-model-list"); + h.listOpenAiModels.mockResolvedValueOnce({ + data: [{ id: "synthetic-openai-model" }], + }); + const { methods } = await import("../server"); + + await expect(methods.listOpenAiModels()).resolves.toEqual({ + data: [{ id: "synthetic-openai-model" }], + }); + + expect(h.openAiOptions).toHaveBeenCalledExactlyOnceWith({ + apiKey: "synthetic-openai-model-list", + }); + expect(h.replicateOptions).not.toHaveBeenCalled(); + expect(h.run).not.toHaveBeenCalled(); + expect(h.getEntitlement).not.toHaveBeenCalled(); + expect(h.ingestUsageEvent).not.toHaveBeenCalled(); + }); +}); diff --git a/editor/lib/ai/__tests__/server.test.ts b/editor/lib/ai/__tests__/server.test.ts index b5a71058ef..17eff5eea3 100644 --- a/editor/lib/ai/__tests__/server.test.ts +++ b/editor/lib/ai/__tests__/server.test.ts @@ -43,7 +43,7 @@ vi.mock("@/lib/auth/organization", () => ({ requireOrganizationId: vi.fn<(...args: never[]) => unknown>(), })); -// Partial passthrough: keep real `catalog`/`gateway`/`modelSpecById`/ +// Partial passthrough: keep real `catalog`/`vercelAiGateway`/`modelSpecById`/ // `tiers`/`byok` (the cost + gate suites depend on them); only stub // `isByokActive` so a test can flip the BYOK carve-out without setting // a module-load env var. @@ -79,7 +79,7 @@ const mockedIsByokActive = vi.mocked(isByokActive); beforeEach(() => { delete process.env.BYOK_OPENROUTER_API_KEY; - delete process.env.BYOK_AI_GATEWAY_API_KEY; + delete process.env.BYOK_VERCEL_AI_GATEWAY_API_KEY; mockedGetEntitlement.mockReset(); mockedIngestUsageEvent.mockReset(); mockedRefreshBalance.mockReset(); @@ -101,13 +101,13 @@ function stubAuthed(orgId = 7) { } afterEach(() => { - delete process.env.REPLICATE_API_TOKEN; + delete process.env.GG_REPLICATE_API_TOKEN; vi.clearAllMocks(); }); describe("methods.generateMusic", () => { it("sends only fields accepted by the current Replicate Lyria schema", async () => { - process.env.REPLICATE_API_TOKEN = "test-replicate-token"; + process.env.GG_REPLICATE_API_TOKEN = "test-replicate-token"; mockedGetEntitlement.mockResolvedValueOnce({ allowed: true, cachedBalanceCents: 1000, diff --git a/editor/lib/ai/image-cost.ts b/editor/lib/ai/image-cost.ts index eb91dbf33d..d14e29d68d 100644 --- a/editor/lib/ai/image-cost.ts +++ b/editor/lib/ai/image-cost.ts @@ -22,7 +22,7 @@ const QUALITY_TIERS = new Set(["high", "medium", "low"]); * `avg_cost_usd` — the documented pricing-anchor behavior for * in-envelope but off-preset sizes. * - `per_token` — fallback to the card's documented average approximation. - * The billing middleware prefers Gateway's actual response cost when present. + * The billing middleware prefers Vercel AI Gateway's actual response cost when present. */ export function computeImageCostMills( card: ai.image.ImageModelCard, diff --git a/editor/lib/ai/models.ts b/editor/lib/ai/models.ts index e883afd03a..7ee69398d1 100644 --- a/editor/lib/ai/models.ts +++ b/editor/lib/ai/models.ts @@ -1,7 +1,7 @@ /** * Editor-side AI provider seam — `GRIDA-SEC-003` carve-out. * - * Owns the attributed AI Gateway provider ({@link gateway}) and the + * Owns the attributed Vercel AI Gateway provider ({@link vercelAiGateway}) and the * BYOK branch ({@link byok}). All catalogue data (text-model specs, * tier→spec map, lookup helpers) lives in `@grida/ai-models/grida` under * its `catalog.text.*` namespace and is re-exported here under its @@ -17,7 +17,7 @@ * @module */ -import { createGateway } from "ai"; +import { createGateway as createVercelAiGateway } from "ai"; import { createOpenAICompatible } from "@ai-sdk/openai-compatible"; import { catalog as _catalog, TIER_MODEL_IDS } from "@grida/ai-models/grida"; @@ -51,24 +51,24 @@ export const models = _catalog.text.byTier; export const tiers = TIER_MODEL_IDS; // --------------------------------------------------------------------------- -// AI Gateway instance — with app attribution headers +// Vercel AI Gateway instance — with app attribution headers // // @see https://vercel.com/docs/ai-gateway/ecosystem/app-attribution // --------------------------------------------------------------------------- /** * Vercel AI Gateway app-attribution headers — lowercase per the `ai` - * SDK convention. Shared by the billed `gateway` and the BYOK - * AI-Gateway branch. (OpenRouter uses its own `HTTP-Referer`/`X-Title` + * SDK convention. Shared by the billed `vercelAiGateway` and the BYOK + * Vercel AI Gateway branch. (OpenRouter uses its own `HTTP-Referer`/`X-Title` * casing — see `resolveByokProvider`.) */ -const GATEWAY_ATTRIBUTION_HEADERS = { +const VERCEL_AI_GATEWAY_ATTRIBUTION_HEADERS = { "http-referer": "https://grida.co", "x-title": "Grida", } as const; /** - * Attributed AI Gateway instance. + * Attributed Vercel AI Gateway instance. * * **Internal — seam consumers only.** This is the raw Vercel AI Gateway * provider; it does NOT go through the billing seam. Any code outside @@ -76,21 +76,26 @@ const GATEWAY_ATTRIBUTION_HEADERS = { * [editor/lib/ai/server.ts](./server.ts) instead, which wraps every model * with gate + ingest middleware. * - * Lint blocks direct imports of this export from non-seam files (see - * [editor/.oxlintrc.jsonc](../../.oxlintrc.jsonc)). + * Keep raw-provider imports inside the seam. The SDK package restrictions in + * [editor/.oxlintrc.jsonc](../../.oxlintrc.jsonc) do not enforce named exports. */ -export const gateway = createGateway({ - headers: GATEWAY_ATTRIBUTION_HEADERS, +// GRIDA-GG: gateway — explicit Grida key or Vercel's platform OIDC authority. +export const vercelAiGateway = createVercelAiGateway({ + // The SDK treats an explicit empty string as no API key and resolves OIDC + // per request. `undefined` would instead read ambient AI_GATEWAY_API_KEY, + // which belongs to the vendor/native-user contract, not funded authority. + apiKey: process.env.GG_VERCEL_AI_GATEWAY_API_KEY?.trim() ?? "", + headers: VERCEL_AI_GATEWAY_ATTRIBUTION_HEADERS, }); // --------------------------------------------------------------------------- // BYOK layer — GRIDA-SEC-003 carve-out (see /SECURITY.md). // // **Internal — seam consumers only.** `byok` holds a live provider API -// key; like `gateway`, consume only from `lib/ai/server.ts`. (No lint +// key; like `vercelAiGateway`, consume only from `lib/ai/server.ts`. (No lint // rule enforces this for the named export — `.oxlintrc.jsonc` restricts -// SDK *packages*, not `gateway`/`byok` imports — convention only, -// mirroring `gateway`.) +// SDK *packages*, not `vercelAiGateway`/`byok` imports — convention only, +// mirroring `vercelAiGateway`.) // // When a contributor sets a BYOK key, calls route through a BARE // provider that bypasses the billing seam entirely (no gate, no @@ -98,11 +103,11 @@ export const gateway = createGateway({ // provider directly — no Grida balance to meter/drain. BYOK bypasses // billing ONLY, never auth (requireOrganizationId still runs). Gated // solely by server-only env vars NEVER set in the hosted product (same -// trust model as OPENAI_API_KEY / REPLICATE_API_TOKEN). Fail-closed: +// trust model as other server-only credentials). Fail-closed: // active only when a key env var is a non-empty string after trim // (whitespace-only secrets fall back to the billed path). // -// Implementations (precedence: OpenRouter first, then Vercel). A +// Implementations (precedence: OpenRouter first, then Vercel AI Gateway). A // third BYOK key is a new branch here — no registry. // --------------------------------------------------------------------------- function resolveByokProvider() { @@ -115,11 +120,11 @@ function resolveByokProvider() { headers: { "HTTP-Referer": "https://grida.co", "X-Title": "Grida" }, }); } - const aiGatewayKey = process.env.BYOK_AI_GATEWAY_API_KEY?.trim(); - if (aiGatewayKey) { - return createGateway({ - apiKey: aiGatewayKey, - headers: GATEWAY_ATTRIBUTION_HEADERS, + const vercelAiGatewayKey = process.env.BYOK_VERCEL_AI_GATEWAY_API_KEY?.trim(); + if (vercelAiGatewayKey) { + return createVercelAiGateway({ + apiKey: vercelAiGatewayKey, + headers: VERCEL_AI_GATEWAY_ATTRIBUTION_HEADERS, }); } return null; diff --git a/editor/lib/ai/openai-compat/hosted-models.test.ts b/editor/lib/ai/openai-compat/hosted-models.test.ts index 5303e90804..ff6a2800eb 100644 --- a/editor/lib/ai/openai-compat/hosted-models.test.ts +++ b/editor/lib/ai/openai-compat/hosted-models.test.ts @@ -125,7 +125,7 @@ describe("hosted catalog", () => { } }); - it("preserves image/video Vercel availability and the pricing-free public shape", () => { + it("preserves image/video Vercel AI Gateway availability and the pricing-free public shape", () => { const list = hostedModelList(); expect( list diff --git a/editor/lib/ai/openai-compat/hosted-models.ts b/editor/lib/ai/openai-compat/hosted-models.ts index d577ff06d9..fe5fe78612 100644 --- a/editor/lib/ai/openai-compat/hosted-models.ts +++ b/editor/lib/ai/openai-compat/hosted-models.ts @@ -11,7 +11,7 @@ * stay callable and are flagged on `/models`; removing a model from the * listed set withdraws its hosted availability. * - image/video: listed cards carrying a `vercel` binding (what the - * seam can serve through the gateway). + * seam can serve through Vercel AI Gateway). * * Deliberately NO pricing in the payload — pricing is a billing-page * concern; exposing per-token USD here invites client-side cost math diff --git a/editor/lib/ai/server.ts b/editor/lib/ai/server.ts index e9110f1475..10967ae4c8 100644 --- a/editor/lib/ai/server.ts +++ b/editor/lib/ai/server.ts @@ -53,7 +53,7 @@ import { import { byok, catalog, - gateway, + vercelAiGateway, isByokActive, modelSpecById, tiers, @@ -431,7 +431,7 @@ const languageModelMiddleware: LanguageModelMiddleware = { const imageModelMiddleware: ImageModelMiddleware = { specificationVersion: "v3", - // Prefer Gateway's response receipt (USD), which accounts for actual quality, + // Prefer Vercel AI Gateway's response receipt (USD), which accounts for actual quality, // dimensions and cache use. Caller pricing remains a fallback for responses // without a receipt; aggregate image token counts cannot price mixed inputs. // https://vercel.com/academy/ai-gateway/ai-gateway-pricing @@ -444,12 +444,12 @@ const imageModelMiddleware: ImageModelMiddleware = { } return withTransaction(ctx, async () => { const result = await doGenerate(); - const gatewayMetadata = result.providerMetadata?.gateway; + const vercelAiGatewayMetadata = result.providerMetadata?.gateway; const rawCost = - gatewayMetadata && - typeof gatewayMetadata === "object" && - "cost" in gatewayMetadata - ? gatewayMetadata.cost + vercelAiGatewayMetadata && + typeof vercelAiGatewayMetadata === "object" && + "cost" in vercelAiGatewayMetadata + ? vercelAiGatewayMetadata.cost : undefined; const reportedUsd = typeof rawCost === "number" @@ -470,7 +470,7 @@ const imageModelMiddleware: ImageModelMiddleware = { }; const wrappedProvider = wrapProvider({ - provider: gateway, + provider: vercelAiGateway, languageModelMiddleware, imageModelMiddleware, }); @@ -524,7 +524,7 @@ export function model(tier: ModelTier) { // Library composes its query embedding in `@/lib/library/embedding`, which // owns the model id + 1536-d truncation + cache and calls this primitive. -const embeddingProvider = byok ?? gateway; +const embeddingProvider = byok ?? vercelAiGateway; /** * UNBILLED text embedding through the attributed provider (BYOK precedence: @@ -556,10 +556,10 @@ let _replicateClient: Replicate | null = null; function getReplicateClient(): Replicate { if (_replicateClient) return _replicateClient; - const token = process.env.REPLICATE_API_TOKEN; + const token = process.env.GG_REPLICATE_API_TOKEN?.trim(); if (!token) { throw new Error( - "REPLICATE_API_TOKEN is not set; cannot run Replicate predictions." + "GG_REPLICATE_API_TOKEN is not set; cannot run Replicate predictions." ); } _replicateClient = new Replicate({ auth: token, useFileOutput: false }); @@ -791,9 +791,9 @@ export namespace methods { } | null { const card = ai.image.findImageModelCard(model); if (!card) return null; - // Call the gateway by the vercel BINDING id, never the canonical card + // Call Vercel AI Gateway by its binding id, never the canonical card // id. They usually coincide, but not always — the Gemini card's - // canonical key is the `-preview` alias while the gateway's current id + // canonical key is the `-preview` alias while Vercel AI Gateway's current id // is the graduated one. Video already resolves this way. const binding = ai.image.binding(card, "vercel"); if (!binding) return null; @@ -802,10 +802,10 @@ export namespace methods { /** * Hosted video generation — gate → `experimental_generateVideo` via - * the RAW gateway → ingest. Explicit `withTransaction` because + * the raw Vercel AI Gateway provider → ingest. Explicit `withTransaction` because * `wrapProvider` has no video middleware (unlike text/image). * - * Billing is **pre-priced by requested duration** against the vercel + * Billing is **pre-priced by requested duration** against the Vercel AI Gateway * binding's `(resolution-label, audio-mode)` per-second rate — the * same pre-computed-cost pattern as every Replicate/image call. * Actual output duration may differ slightly (bounded by the card's @@ -894,7 +894,7 @@ export namespace methods { { organizationId, feature: "v1/ai/video", model_id: card.id }, async () => { const generation = await experimental_generateVideo({ - model: gateway.videoModel(binding.id), + model: vercelAiGateway.videoModel(binding.id), prompt: req.prompt, aspectRatio: aspect_ratio as `${number}:${number}`, resolution: req.resolution as `${number}x${number}` | undefined, diff --git a/editor/lib/api/README.md b/editor/lib/api/README.md index dd83684c0e..cef3ef5fab 100644 --- a/editor/lib/api/README.md +++ b/editor/lib/api/README.md @@ -189,13 +189,11 @@ a free result or an invented estimate. See Configure server-only `GG_TRIPO_API_KEY`, GG signing, ordinary billing and request rate limiting before deploying. There is no fallback to a development BYOK key. -The `GG_` prefix identifies Grida-funded infrastructure credentials. Tripo is -the first provider adopting this convention. This unreleased integration uses -a direct cutover from `TRIPO_API_KEY`, with no unprefixed or BYOK alias; update -the infrastructure environment variable name before deployment. Native CLI -`TRIPO_API_KEY` remains the user's direct-provider credential and is unchanged. -Remaining provider adoption is tracked in -[the GG credential naming follow-up](https://github.com/gridaco/grida/issues/1066). +The `GG_` prefix identifies Grida-funded infrastructure credentials. Tripo uses +only `GG_TRIPO_API_KEY`, with no unprefixed or BYOK alias. Native CLI +`TRIPO_API_KEY` remains the user's direct-provider credential. See the +[funded credential mapping and deployment migration](https://grida.co/docs/contributing/billing#server-provider-credentials) +for other providers and the Vercel platform OIDC exception. ElevenLabs is unchanged. Desktop advertises the funded paths separately through `tripo_gg` and `rigging_gg`, so older hosts continue to offer only their supported BYOK paths. diff --git a/editor/lib/desktop/bridge.ts b/editor/lib/desktop/bridge.ts index 20e682f9a0..ae599c1eeb 100644 --- a/editor/lib/desktop/bridge.ts +++ b/editor/lib/desktop/bridge.ts @@ -837,7 +837,7 @@ export namespace app { /** * The Grida agent runs behind the AgentHost, - * against BYOK (OpenRouter / AI Gateway) directly. The renderer sends a + * against BYOK (OpenRouter / Vercel AI Gateway) directly. The renderer sends a * `messages` payload (plus optional `workspaceId` / `skills`), * receives an AI SDK UI-message stream, and either resolves fs * tool calls locally (standalone document window: live `SvgEditor` diff --git a/editor/scaffolds/desktop/shared/media-model-availability.test.ts b/editor/scaffolds/desktop/shared/media-model-availability.test.ts index 0036e608c9..627c8d9757 100644 --- a/editor/scaffolds/desktop/shared/media-model-availability.test.ts +++ b/editor/scaffolds/desktop/shared/media-model-availability.test.ts @@ -119,7 +119,7 @@ describe("MediaModelAvailability.image", () => { ).toBe(true); }); - it("only offers hosted readiness on models with a Vercel binding", () => { + it("only offers hosted readiness on models with a Vercel AI Gateway binding", () => { const hosted = { ...ready, configured: [], hosted: true }; expect(MediaModelAvailability.image(gpt2, hosted).available).toBe(true); expect(MediaModelAvailability.image(flare, hosted).available).toBe(false); @@ -150,7 +150,7 @@ describe("MediaModelAvailability.image", () => { ).toBe(true); }); - it("permits hosted transparency only when the Vercel binding verifies it", () => { + it("permits hosted transparency only when the Vercel AI Gateway binding verifies it", () => { const hosted = { ...ready, configured: [], hosted: true }; const verified: models.image.ImageModelCard = { ...gpt2, diff --git a/editor/scaffolds/desktop/shared/media-model-availability.ts b/editor/scaffolds/desktop/shared/media-model-availability.ts index fe2bbc91e1..9ec5ed6f33 100644 --- a/editor/scaffolds/desktop/shared/media-model-availability.ts +++ b/editor/scaffolds/desktop/shared/media-model-availability.ts @@ -95,7 +95,7 @@ export namespace MediaModelAvailability { ) { return { available: true }; } - // GRIDA-GG: desktop — hosted readiness follows the served Vercel binding. + // GRIDA-GG: desktop — hosted readiness follows the served Vercel AI Gateway binding. if ( state.hosted && models.image.binding(card, "vercel") && diff --git a/editor/scaffolds/desktop/video-gen/video-playground.tsx b/editor/scaffolds/desktop/video-gen/video-playground.tsx index 2f0b3fef3b..645dfd3aa5 100644 --- a/editor/scaffolds/desktop/video-gen/video-playground.tsx +++ b/editor/scaffolds/desktop/video-gen/video-playground.tsx @@ -120,8 +120,8 @@ type Tile = { * floating prompt composer. Each submit prepends a shimmering cell that fills in * with a playable clip when the video resolves. Generation runs in the agent * sidecar against the user's connected provider key; the key never reaches this - * renderer (GRIDA-SEC-004). v1 is text-to-video (served by a connected Vercel - * key for every listed model). + * renderer (GRIDA-SEC-004). v1 is text-to-video (served by a connected + * Vercel AI Gateway key for every listed model). */ export function DesktopVideoPlayground({ initialModelId, diff --git a/editor/turbo.json b/editor/turbo.json index bf42dd7e97..10ffe66143 100644 --- a/editor/turbo.json +++ b/editor/turbo.json @@ -17,6 +17,8 @@ "BIRD_WORKSPACE_ID", "EDGE_CONFIG", "ENABLE_EXPERIMENTAL_COREPACK", + "GG_REPLICATE_API_TOKEN", + "GG_VERCEL_AI_GATEWAY_API_KEY", "GRIDA_S2S_PRIVATE_API_KEY", "INTEGRATIONS_TEST_TOSSPAYMENTS_CUSTOMER_KEY", "INTEGRATIONS_TEST_TOSSPAYMENTS_SECRET_KEY", @@ -30,7 +32,6 @@ "POSTGRES_URL_NON_POOLING", "POSTGRES_USER", "REDIS_URL", - "REPLICATE_API_TOKEN", "RESEND_API_KEY", "SENTRY_AUTH_TOKEN", "SENTRY_HOST", diff --git a/packages/grida-ai-agent/README.md b/packages/grida-ai-agent/README.md index 1eda192405..adf804f15f 100644 --- a/packages/grida-ai-agent/README.md +++ b/packages/grida-ai-agent/README.md @@ -229,11 +229,11 @@ rejection is terminal; there is no ambient-download fallback. Provider result URLs are narrower than the callback's general authorization surface. OpenRouter video ignores third-party `unsigned_urls` and fetches only -its authenticated, same-origin content endpoint. Vercel Gateway video accepts -only inline `data:` results or its exact configured Gateway origin; an +its authenticated, same-origin content endpoint. Vercel AI Gateway video accepts +only inline `data:` results or its exact configured Vercel AI Gateway origin; an arbitrary result origin fails with `unsupported_untrusted_result_origin`, and -an exact remote origin still requires the host download transport. Vercel -Gateway image responses are base64 strings on the provider request lane and +an exact remote origin still requires the host download transport. Vercel AI Gateway image +responses are base64 strings on the provider request lane and never open the download lane. Image generation accepts optional native `background` intent (`auto`, `opaque`, @@ -247,10 +247,11 @@ provider options. Generated bytes are saved unchanged, preserving alpha. Hosted requests enforce the same admission locally and forward the requirement for independent server-side validation before billing. -GPT Image 2.5 Flare and Sunburst bind Vercel, OpenRouter, and fal. Vercel and fal +GPT Image 2.5 Flare and Sunburst bind Vercel AI Gateway, OpenRouter, and fal. Vercel AI Gateway +and fal have verified native-background controls; OpenRouter's published endpoint schema does not expose transparency. Reference-conditioned generation is available on -OpenRouter and fal; transparent edits require fal because Vercel reference +OpenRouter and fal; transparent edits require fal because Vercel AI Gateway reference bindings are not yet verified. OpenRouter uses the same canonical model id for generation and edits, while fal uses separate `/text-to-image` and `/edit` routes. All three accept the six quality levels through their own provider namespace. @@ -300,7 +301,7 @@ crosses one of these is the wrong tool, not a missing feature. - **Not a general model-provider router.** Provider selection is isolated to the node-only `providers/` layer: the BYOK key slots - (OpenRouter → AI Gateway) plus ONE generalized OpenAI-compatible + (OpenRouter → Vercel AI Gateway) plus ONE generalized OpenAI-compatible endpoint type (`{base_url, optional key, registered models}` — Ollama is the preset; issue #806), and the narrowly configured native ChatGPT subscription provider. The agent + runtime core never import diff --git a/packages/grida-ai-agent/src/__public-api__.test.ts b/packages/grida-ai-agent/src/__public-api__.test.ts index d710cbcce7..7a7196e0dd 100644 --- a/packages/grida-ai-agent/src/__public-api__.test.ts +++ b/packages/grida-ai-agent/src/__public-api__.test.ts @@ -138,7 +138,7 @@ describe("@grida/agent public API", () => { ]); expect(BYOK_PROVIDER_METADATA.map((provider) => provider.label)).toEqual([ "OpenRouter", - "Vercel", + "Vercel AI Gateway", "fal", "ElevenLabs", "Tripo", diff --git a/packages/grida-ai-agent/src/http/routes/video.test.ts b/packages/grida-ai-agent/src/http/routes/video.test.ts index 49e30483c4..65e5a183b2 100644 --- a/packages/grida-ai-agent/src/http/routes/video.test.ts +++ b/packages/grida-ai-agent/src/http/routes/video.test.ts @@ -40,13 +40,13 @@ function post(app: Hono, payload: unknown, signal?: AbortSignal) { }); } -// Veo binds vercel + fal; fal is a plain-fetch adapter we can drive end-to-end. +// Veo binds Vercel AI Gateway + fal; fal uses the plain-fetch adapter. const VEO = "google/veo-3.1"; /** Shape of the `init` arg our `fetch` mock reads. */ type MockInit = { method?: string; body?: string }; -function vercelVideoResults(urls: string[]): Response { +function vercelAiGatewayVideoResults(urls: string[]): Response { return new Response( `data: ${JSON.stringify({ type: "result", @@ -64,8 +64,8 @@ function vercelVideoResults(urls: string[]): Response { ); } -function vercelVideoResult(url: string): Response { - return vercelVideoResults([url]); +function vercelAiGatewayVideoResult(url: string): Response { + return vercelAiGatewayVideoResults([url]); } afterEach(() => vi.unstubAllGlobals()); @@ -150,12 +150,14 @@ describe("POST /video/generate", () => { expect(download).toHaveBeenCalledOnce(); }); - it("rejects an arbitrary Vercel result origin before host download", async () => { + it("rejects an arbitrary Vercel AI Gateway result origin before host download", async () => { const request = vi.fn(async (input) => { expect(String(input)).toBe( "https://ai-gateway.vercel.sh/v3/ai/video-model" ); - return vercelVideoResult("https://vendor-cdn.example/video.mp4?token=x"); + return vercelAiGatewayVideoResult( + "https://vendor-cdn.example/video.mp4?token=x" + ); }); const download = vi.fn(); const providerHttp = new ProviderHttp({ request, download }); @@ -175,9 +177,9 @@ describe("POST /video/generate", () => { expect(download).not.toHaveBeenCalled(); }); - it("decodes an inline Vercel data result without host download", async () => { + it("decodes an inline Vercel AI Gateway data result without host download", async () => { const request = vi.fn(async () => - vercelVideoResult("data:video/mp4;base64,AAAY") + vercelAiGatewayVideoResult("data:video/mp4;base64,AAAY") ); const download = vi.fn(); const providerHttp = new ProviderHttp({ request, download }); @@ -196,11 +198,11 @@ describe("POST /video/generate", () => { expect(download).not.toHaveBeenCalled(); }); - it("permits an exact Vercel Gateway result origin through host download", async () => { + it("permits an exact Vercel AI Gateway result origin through host download", async () => { const MP4 = new Uint8Array([0, 0, 0, 24]); const resultUrl = "https://ai-gateway.vercel.sh/results/video.mp4"; const request = vi.fn(async () => - vercelVideoResult(resultUrl) + vercelAiGatewayVideoResult(resultUrl) ); const download = vi.fn(async (input) => { expect(String(input)).toBe(resultUrl); @@ -234,7 +236,7 @@ describe("POST /video/generate", () => { const token = "opaque-signed-capability"; const resultUrl = `https://ai-gateway.vercel.sh/results/video.mp4?X-Amz-Signature=${token}`; const request = vi.fn(async () => - vercelVideoResult(resultUrl) + vercelAiGatewayVideoResult(resultUrl) ); const download = vi.fn( async () => @@ -256,10 +258,10 @@ describe("POST /video/generate", () => { expect(download).toHaveBeenCalledOnce(); }); - it("fails closed on a remote Vercel result without host download authority", async () => { + it("fails closed on a remote Vercel AI Gateway result without host download authority", async () => { const resultUrl = "https://ai-gateway.vercel.sh/v3/ai/video-result.mp4"; const ambient = vi.fn<(input: string | URL | Request) => Promise>( - async () => vercelVideoResult(resultUrl) + async () => vercelAiGatewayVideoResult(resultUrl) ); vi.stubGlobal("fetch", ambient); @@ -338,7 +340,7 @@ describe("POST /video/generate", () => { ); it.each([{ duration: -1 }, { seed: 0 }])( - "rejects unsupported Vercel options before provider I/O or receipts: %j", + "rejects unsupported Vercel AI Gateway options before provider I/O or receipts: %j", async (options) => { const request = vi.fn(); const save = vi.fn(); @@ -526,7 +528,7 @@ describe("POST /video/generate", () => { }); it("400 when the connected provider does not serve the model", async () => { - // Seedance has no vercel binding (token-metered on the gateway; the + // Seedance has no Vercel AI Gateway binding (token-metered there; the // catalogue withholds an unpriceable route). fal DOES serve it — which // is why this case must not use a fal key: it would reach fal for real. const res = await post(appWith({ vercel: "sk-v" }), { @@ -553,7 +555,7 @@ describe("POST /video/generate", () => { it("offers every output to the host store and correlates only accepted descriptors", async () => { const request = vi.fn(async () => - vercelVideoResults([ + vercelAiGatewayVideoResults([ "data:video/mp4;base64,AAAY", "data:video/mp4;base64,AQID", ]) diff --git a/packages/grida-ai-agent/src/providers/byok.ts b/packages/grida-ai-agent/src/providers/byok.ts index b0fd46a502..504aa2faac 100644 --- a/packages/grida-ai-agent/src/providers/byok.ts +++ b/packages/grida-ai-agent/src/providers/byok.ts @@ -6,7 +6,7 @@ */ import { createOpenAICompatible } from "@ai-sdk/openai-compatible"; -import { createGateway } from "@ai-sdk/gateway"; +import { createGateway as createVercelAiGateway } from "@ai-sdk/gateway"; import { TIER_MODEL_IDS, type TierModelId } from "@grida/ai-models/grida"; import type { ModelFactory } from "../agent"; import type { ModelTier } from "../tiers"; @@ -47,18 +47,21 @@ export function makeOpenRouterFactory( includeUsage: true, fetch: providerHttp.request, }); - // Both OpenRouter and the catalog use Vercel-style `creator/model` + // Both OpenRouter and the catalog use the `creator/model` format for // ids, so an explicit pick hands straight through; otherwise fall // back to the tier's canonical model. return (tier, modelId) => provider(modelId ?? tierModelIds()[tier]); } -export function makeVercelFactory( +export function makeVercelAiGatewayFactory( apiKey: string, providerHttp: ProviderHttp = new ProviderHttp(), tierModelIds: TierModelIds = BUNDLED_TIER_MODEL_IDS ): ModelFactory { - const provider = createGateway({ apiKey, fetch: providerHttp.request }); + const provider = createVercelAiGateway({ + apiKey, + fetch: providerHttp.request, + }); return (tier, modelId) => provider(modelId ?? tierModelIds()[tier]); } diff --git a/packages/grida-ai-agent/src/providers/http.test.ts b/packages/grida-ai-agent/src/providers/http.test.ts index 5a3b15fa0c..38089c154a 100644 --- a/packages/grida-ai-agent/src/providers/http.test.ts +++ b/packages/grida-ai-agent/src/providers/http.test.ts @@ -3,7 +3,7 @@ import { describe, expect, it, vi } from "vitest"; import { makeEndpointFactory, makeOpenRouterFactory, - makeVercelFactory, + makeVercelAiGatewayFactory, } from "./byok"; import { ProviderHttp } from "./http"; import { probeEndpointModels } from "./probe"; @@ -22,7 +22,7 @@ async function attemptModelRequest(model: unknown): Promise { describe("ProviderHttp", () => { it("every in-process text provider uses request, never download", async () => { // The thrown sentinel stops each real SDK immediately after it crosses the - // transport seam; URL assertions prove OpenRouter, Vercel Gateway, and the + // transport seam; URL assertions prove OpenRouter, Vercel AI Gateway, and the // generalized OpenAI-compatible endpoint all reached the supplied request. const urls: string[] = []; const request: typeof globalThis.fetch = async (input) => { @@ -35,7 +35,7 @@ describe("ProviderHttp", () => { const http = new ProviderHttp({ request, download }); await attemptModelRequest(makeOpenRouterFactory("sk-or", http)("pro")); - await attemptModelRequest(makeVercelFactory("sk-v", http)("pro")); + await attemptModelRequest(makeVercelAiGatewayFactory("sk-v", http)("pro")); await attemptModelRequest( makeEndpointFactory( { @@ -186,7 +186,7 @@ describe("ProviderHttp", () => { expect(download).not.toHaveBeenCalled(); }); - it("OpenRouter and Vercel media SDK calls use request", async () => { + it("OpenRouter and Vercel AI Gateway media SDK calls use request", async () => { const urls: string[] = []; const request: typeof globalThis.fetch = async (input) => { urls.push(String(input)); diff --git a/packages/grida-ai-agent/src/providers/index.test.ts b/packages/grida-ai-agent/src/providers/index.test.ts index 3f8c412995..91e92c5d0a 100644 --- a/packages/grida-ai-agent/src/providers/index.test.ts +++ b/packages/grida-ai-agent/src/providers/index.test.ts @@ -30,7 +30,7 @@ const OLLAMA: EndpointProviderConfig = { }; describe("resolveProvider", () => { - it("prefers OpenRouter over Vercel when both BYOK keys exist", async () => { + it("prefers OpenRouter over Vercel AI Gateway when both BYOK keys exist", async () => { const provider = await resolveProvider( deps({ openrouter: " sk-or ", @@ -43,7 +43,7 @@ describe("resolveProvider", () => { expect(provider.model_factory).toBeTypeOf("function"); }); - it("falls back to Vercel when OpenRouter is absent", async () => { + it("falls back to Vercel AI Gateway when OpenRouter is absent", async () => { const provider = await resolveProvider(deps({ vercel: "vercel-key" })); expect(provider.provider_id).toBe("vercel"); diff --git a/packages/grida-ai-agent/src/providers/index.ts b/packages/grida-ai-agent/src/providers/index.ts index 1a421ab508..63881f7732 100644 --- a/packages/grida-ai-agent/src/providers/index.ts +++ b/packages/grida-ai-agent/src/providers/index.ts @@ -13,7 +13,7 @@ * * - `chatgpt` — a native ChatGPT subscription credential whose model * adapter still runs inside Grida's agent loop. - * - `byok` — the hardcoded third-party slots (OpenRouter, Vercel), + * - `byok` — the hardcoded third-party slots (OpenRouter, Vercel AI Gateway), * keyed by a stored secret. * - `gg` — Grida's hosted, included provider. * - `endpoint` — ONE generalized OpenAI-compatible endpoint type @@ -53,7 +53,7 @@ import type { EndpointProvidersStore } from "./endpoints"; import { makeEndpointFactory, makeOpenRouterFactory, - makeVercelFactory, + makeVercelAiGatewayFactory, } from "./byok"; import type { ProviderHttp } from "./http"; import { ChatGptProvider, type ChatGptProviderRuntime } from "./chatgpt"; @@ -324,7 +324,7 @@ function makeResolvedByok( return { provider_id: providerId, kind: "byok", - model_factory: makeVercelFactory( + model_factory: makeVercelAiGatewayFactory( key.trim(), providerHttp, tierModelIds diff --git a/packages/grida-ai-agent/src/providers/resolve-image.test.ts b/packages/grida-ai-agent/src/providers/resolve-image.test.ts index eb42a5bc8f..5553107d4c 100644 --- a/packages/grida-ai-agent/src/providers/resolve-image.test.ts +++ b/packages/grida-ai-agent/src/providers/resolve-image.test.ts @@ -18,7 +18,7 @@ function fakeSecrets(keys: Record): SecretsStore { } as unknown as SecretsStore; } -// A universal listed card (binds vercel + fal + openrouter). +// A universal listed card (binds Vercel AI Gateway + fal + OpenRouter). const LISTED = "openai/gpt-image-2"; // A kept-but-unlisted card (bfl/flux-kontext-max — not on OpenRouter). const UNLISTED = "bfl/flux-kontext-max"; @@ -57,7 +57,7 @@ describe("defaultImageModelId", () => { describe("resolveImageModel", () => { describe("native background requirements", () => { it.each(["opaque", "transparent"] as const)( - "selects verified FAL for %s before OpenRouter or Vercel", + "selects verified FAL for %s before OpenRouter or Vercel AI Gateway", async (background) => { const readKey = vi.fn( async (id) => `key-${id}` @@ -338,7 +338,7 @@ describe("resolveImageModel", () => { } ); - it("selects Vercel for transparency before FAL while skipping OpenRouter", async () => { + it("selects Vercel AI Gateway for transparency before FAL while skipping OpenRouter", async () => { const r = await resolveImageModel( { secrets: fakeSecrets({ @@ -425,7 +425,7 @@ describe("resolveImageModel", () => { } ); - it("withholds unverified Vercel and GG edit routes", async () => { + it("withholds unverified Vercel AI Gateway and GG edit routes", async () => { await expect( resolveImageModel({ secrets: fakeSecrets({ vercel: "sk-v" }) }, id, { references: true, diff --git a/packages/grida-ai-agent/src/providers/resolve-video.test.ts b/packages/grida-ai-agent/src/providers/resolve-video.test.ts index e9539c9d4e..c19b9b5d33 100644 --- a/packages/grida-ai-agent/src/providers/resolve-video.test.ts +++ b/packages/grida-ai-agent/src/providers/resolve-video.test.ts @@ -8,8 +8,8 @@ function fakeSecrets(keys: Record): SecretsStore { } as unknown as SecretsStore; } -// Veo 3.1 binds vercel + fal; Seedance 2.0 binds fal + openrouter but NOT -// vercel — the gateway meters it per token, which the catalogue cannot price. +// Veo 3.1 binds Vercel AI Gateway + fal; Seedance 2.0 binds fal + OpenRouter. +// Vercel AI Gateway meters Seedance per token, which the catalogue cannot price. const VEO = "google/veo-3.1"; const SEEDANCE = "bytedance/seedance-2.0"; @@ -27,7 +27,7 @@ describe("resolveVideoModel", () => { expect(Object.isFrozen(r)).toBe(true); }); - it("prefers Vercel over fal when both keys exist", async () => { + it("prefers Vercel AI Gateway over fal when both keys exist", async () => { const r = await resolveVideoModel( { secrets: fakeSecrets({ vercel: "sk-v", fal: "sk-fal" }) }, VEO @@ -37,7 +37,7 @@ describe("resolveVideoModel", () => { }); it("falls through when the only key's provider does not serve the model", async () => { - // Seedance has no vercel binding — a Vercel-only user can't run it. + // Seedance has no Vercel AI Gateway binding — a Vercel AI Gateway-only user can't run it. await expect( resolveVideoModel({ secrets: fakeSecrets({ vercel: "sk-v" }) }, SEEDANCE) ).rejects.toBeInstanceOf(VideoModelUnavailableError); diff --git a/packages/grida-ai-agent/src/runtime/image-generation.test.ts b/packages/grida-ai-agent/src/runtime/image-generation.test.ts index 79d9f202d3..fc87a89ddb 100644 --- a/packages/grida-ai-agent/src/runtime/image-generation.test.ts +++ b/packages/grida-ai-agent/src/runtime/image-generation.test.ts @@ -99,7 +99,7 @@ describe("workspace image operation consumer", () => { it("resolves reference support before any workspace bytes are read", async () => { const { bindings, generate, request } = await build({ - // The default's fal route now supports references; Vercel's does not. + // The default's fal route now supports references; Vercel AI Gateway's does not. keys: { vercel: "synthetic-key" }, }); const read = vi.spyOn(bindings.fs, "readBytes"); diff --git a/packages/grida-ai-models/README.md b/packages/grida-ai-models/README.md index d01c710ce5..6a703cdbc9 100644 --- a/packages/grida-ai-models/README.md +++ b/packages/grida-ai-models/README.md @@ -70,7 +70,7 @@ omitted declaration inherits the model. Missing bindings never grant support. Image pricing may be tiered, flat, or per-token. Published model-level prices remain distinct from provider-binding prices. GPT Image 2's model-level table, for example, includes rectangular image tiers and token components absent from -its Vercel binding. No primary-provider preference is stored on factual cards. +its Vercel AI Gateway binding. No primary-provider preference is stored on factual cards. Video cards identify canonical models and exact image-to-video provider routes. Each binding has its own resolution/audio price matrix and any input-image diff --git a/packages/grida-ai-models/__tests__/grida/modelSpecById.test.ts b/packages/grida-ai-models/__tests__/grida/modelSpecById.test.ts index 6f5f8c16f8..9dbc81c3b9 100644 --- a/packages/grida-ai-models/__tests__/grida/modelSpecById.test.ts +++ b/packages/grida-ai-models/__tests__/grida/modelSpecById.test.ts @@ -26,7 +26,7 @@ describe("models.text.modelSpecById", () => { id: "google/gemini-3.8-flash", label: "Gemini 3.8 Flash", }, - ])("resolves the exact $id gateway id", ({ id, label }) => { + ])("resolves the exact $id provider model id", ({ id, label }) => { const spec = models.text.modelSpecById(id); expect(spec?.id).toBe(id); expect(spec?.label).toBe(label); diff --git a/packages/grida-ai-models/__tests__/grida/models.test.ts b/packages/grida-ai-models/__tests__/grida/models.test.ts index fe4259980c..f87416feaa 100644 --- a/packages/grida-ai-models/__tests__/grida/models.test.ts +++ b/packages/grida-ai-models/__tests__/grida/models.test.ts @@ -66,7 +66,7 @@ describe("bundled model release metadata", () => { }); describe("models.image.findImageModelCard", () => { - it("resolves a full vercel id", () => { + it("resolves a full Vercel AI Gateway id", () => { const card = models.image.findImageModelCard("bfl/flux-pro-1.1"); expect(card?.id).toBe("bfl/flux-pro-1.1"); expect(card?.label).toBe("Flux Pro 1.1"); @@ -186,16 +186,16 @@ describe("models.image provider-binding invariants", () => { "openrouter", "vercel", ]); - const vercel = models.image.binding(card, "vercel")!; - expect(vercel.id).toBe(card.id); - expect(vercel.pricing).toEqual({ + const vercelAiGatewayBinding = models.image.binding(card, "vercel")!; + expect(vercelAiGatewayBinding.id).toBe(card.id); + expect(vercelAiGatewayBinding.pricing).toEqual({ type: "per_token", input: 5, cached_input: 1.25, output: 30, }); - expect(card.pricing).toEqual(vercel.pricing); - expect(vercel.references).toBeUndefined(); + expect(card.pricing).toEqual(vercelAiGatewayBinding.pricing); + expect(vercelAiGatewayBinding.references).toBeUndefined(); expect(models.image.supportsTransparentBackground(card, "vercel")).toBe( true ); @@ -240,7 +240,7 @@ describe("models.image provider-binding invariants", () => { }); it("prices Recraft V4.1 at the raster rate every provider publishes", () => { - // Written from the providers, not the card: Vercel feed `pricing.image`, + // Written from the providers, not the card: Vercel AI Gateway feed `pricing.image`, // OpenRouter `/images/models/.../endpoints` `cost_usd`, and fal's model // page payload all say $0.035 (checked 2026-09-02). V3 is $0.04 on the // same three, which is why V4.1 is listed and V3 is not. @@ -260,9 +260,9 @@ describe("models.image provider-binding invariants", () => { }); it("prices Flux Kontext Pro at $0.04 and binds fal", () => { - // The card shipped at $0.05 while the gateway feed's `pricing.image` is + // The card shipped at $0.05 while the Vercel AI Gateway feed's `pricing.image` is // $0.04 — reachable over-deduction, since the hosted route gates on a - // vercel binding rather than `listed`. fal's page states the same $0.04. + // Vercel AI Gateway binding rather than `listed`. fal's page states the same $0.04. const kontext = models.image.models["bfl/flux-kontext-pro"]!; expect(kontext.pricing).toEqual({ type: "per_image_flat", usd: 0.04 }); expect(models.image.binding(kontext, "vercel")?.pricing).toEqual({ @@ -276,8 +276,8 @@ describe("models.image provider-binding invariants", () => { }); it("prices Flux 2 Pro and Max at the per-megapixel baseline every provider publishes", () => { - // Vercel model page, OpenRouter `cost_usd`/megapixel, fal "first megapixel" - // — all $0.03 (Pro) and $0.07 (Max), 2026-09-02. Pro's Vercel binding had + // Vercel AI Gateway model page, OpenRouter `cost_usd`/megapixel, fal "first megapixel" + // — all $0.03 (Pro) and $0.07 (Max), 2026-09-02. Pro's Vercel AI Gateway binding had // shipped at $0.06: a 2x hosted over-billing. for (const [id, usd] of [ ["bfl/flux-2-pro", 0.03], @@ -296,9 +296,9 @@ describe("models.image provider-binding invariants", () => { }); it("lists Seedream 5.0 and deprecates 4.5", () => { - // Vercel feed `pricing.image` $0.035 for both 5.0 cards; fal $0.035 (Lite); + // Vercel AI Gateway feed `pricing.image` $0.035 for both 5.0 cards; fal $0.035 (Lite); // OpenRouter `cost_usd` $0.035 (Lite) / $0.045 (Pro 1K). 4.5 is $0.04 on - // Vercel and fal — dominated by Lite on price and generation. + // Vercel AI Gateway and fal — dominated by Lite on price and generation. const lite = models.image.models["bytedance/seedream-5.0-lite"]!; expect(lite.listed).toBe(true); for (const p of ["vercel", "fal", "openrouter"] as const) { @@ -457,7 +457,7 @@ describe("models.image.supportsTransparentBackground", () => { it("declares only verified native transparency and provider exposure", () => { // Published fal schemas expose transparent backgrounds for GPT Image 2 - // and both 2.5 variants. OpenRouter's enums exclude it. Vercel publishes + // and both 2.5 variants. OpenRouter's enums exclude it. Vercel AI Gateway publishes // support for 2.5 but not GPT Image 2 (verified 2026-09-09 KST). for (const id of [ "openai/gpt-image-2", @@ -725,12 +725,12 @@ describe("models.video catalogue invariants", () => { } }); - it("prices Veo 3.1 with the gateway's full audio/resolution matrix", () => { + it("prices Veo 3.1 with Vercel AI Gateway's full audio/resolution matrix", () => { // Written from the provider, not from the card: these are the exact - // `video_duration_pricing` rows the Vercel gateway's /v1/models feed + // `video_duration_pricing` rows the Vercel AI Gateway's /v1/models feed // returns for `google/veo-3.1-generate-001` (checked 2026-09-02). // The card previously claimed audio-on-only at <=1080p, so a 4K request - // was rejected for want of a rate. The gateway meters both modes, and + // was rejected for want of a rate. Vercel AI Gateway meters both modes, and // fal meters the same matrix. // https://vercel.com/ai-gateway/models/veo-3.1-generate-001 const matrix = { @@ -746,7 +746,7 @@ describe("models.video catalogue invariants", () => { } }); - it("prices Veo 3.1 Fast and Lite with the matrices Vercel and fal both publish", () => { + it("prices Veo 3.1 Fast and Lite with the matrices Vercel AI Gateway and fal both publish", () => { const fast = { "720p": { audio: 0.15, silent: 0.1 }, "1080p": { audio: 0.15, silent: 0.1 }, @@ -769,7 +769,7 @@ describe("models.video catalogue invariants", () => { } }); - it("prices Wan 3.0 per resolution, audio bundled, identically on Vercel and fal", () => { + it("prices Wan 3.0 per resolution, audio bundled, identically on Vercel AI Gateway and fal", () => { const wan = models.video.models["alibaba/wan-3.0"]!; expect(wan).toMatchObject({ min_duration: 2, max_duration: 30 }); for (const p of ["vercel", "fal"] as const) { @@ -781,10 +781,10 @@ describe("models.video catalogue invariants", () => { } }); - it("withholds a Vercel binding from token-metered Seedance", () => { - // The gateway serves both Seedance cards but bills `video_token_pricing`, + it("withholds a Vercel AI Gateway binding from token-metered Seedance", () => { + // Vercel AI Gateway serves both Seedance cards but bills `video_token_pricing`, // which `PerSecondPricing` cannot express and the hosted route cannot - // pre-price. A per-second Vercel binding here would be invented — and + // pre-price. A per-second Vercel AI Gateway binding here would be invented — and // was, on 2.0, until 2026-09-02. fal states per-second rates. for (const id of ["bytedance/seedance-2.0", "bytedance/seedance-2.5"]) { const card = models.video.models[id]!; @@ -811,7 +811,7 @@ describe("models.video catalogue invariants", () => { const grok = models.video.models["xai/grok-imagine-video-1.5"]!; expect(grok).toMatchObject({ min_duration: 1, max_duration: 15 }); - // Vercel serves every SpaceXAI model under `spacexai/`, so the call id is not + // Vercel AI Gateway serves every SpaceXAI model under `spacexai/`, so the call id is not // this card's canonical `xai/` id. Reusing the canonical id here is a 404 // at call time — which is how the binding was first shipped. // https://vercel.com/ai-gateway/models/grok-imagine-video-1.5 diff --git a/packages/grida-ai-models/src/grida/catalog.ts b/packages/grida-ai-models/src/grida/catalog.ts index 3af2160bc7..b6de2957bf 100644 --- a/packages/grida-ai-models/src/grida/catalog.ts +++ b/packages/grida-ai-models/src/grida/catalog.ts @@ -1633,7 +1633,7 @@ export namespace catalog { * UNKNOWN PROVIDER KEYS ARE DROPPED rather than rejecting the card. * A provider this client has no adapter for is not an error — it is a * route it cannot take — and rejecting would make adding a provider a - * breaking publish. (The AI SDK gateway learned this the hard way: it + * breaking publish. (Vercel AI Gateway learned this the hard way: it * validated its model-kind field as a hard enum, so the day a new kind * shipped the whole listing failed to parse; it now accepts loosely * and filters unknown rows.) @@ -1847,7 +1847,7 @@ export namespace catalog { if (!providers) return undefined; // A listed model can launch on one provider first. Its primary route // must be known and bound; runtime selection intersects the available - // bindings with connected keys, and hosted calls still require Vercel. + // bindings with connected keys, and hosted calls still require Vercel AI Gateway. if (!image.providers.includes(v.provider as image.ImageProvider)) { return undefined; } diff --git a/packages/grida-ai-models/src/models.ts b/packages/grida-ai-models/src/models.ts index 57f5023a70..c4f74f3056 100644 --- a/packages/grida-ai-models/src/models.ts +++ b/packages/grida-ai-models/src/models.ts @@ -353,8 +353,8 @@ export namespace models { }, }, // OpenAI's September 3 introduction remained a limited rollout. The - // September 4 date below is the exact Vercel route's broad availability, - // which is the release fact relevant to this gateway-shaped card. + // September 4 date below is the exact Vercel AI Gateway route's broad availability, + // which is the release fact relevant to this route card. // https://openai.com/products/release-notes/ // https://vercel.com/ai-gateway/models/gpt-6-astra "openai/gpt-6-astra": { @@ -855,7 +855,7 @@ export namespace models { /** * A provider that can serve an image model. Distinct from the top-level * {@link models.Provider} because the same flagship proprietary model is - * now multi-homed: fal, OpenRouter, and the Vercel gateway each serve it + * now multi-homed: fal, OpenRouter, and the Vercel AI Gateway each serve it * under a different id (and sometimes a different meter). Mirrors * {@link video.VideoProvider}. */ @@ -870,7 +870,7 @@ export namespace models { export type ImageProviderBinding = { provider: ImageProvider; /** - * Provider-specific call id, e.g. `openai/gpt-image-2` (Vercel) or + * Provider-specific call id, e.g. `openai/gpt-image-2` (Vercel AI Gateway) or * `openai/gpt-image-2.5/flare/text-to-image` (fal). */ id: string; @@ -1026,7 +1026,7 @@ export namespace models { // Each serving provider publishes its own meter; do not copy fal's extra // image-cache/text-output rates into providers that do not advertise them. // https://ai-gateway.vercel.sh/v1/models (verified 2026-09-09 KST) - const GPT_IMAGE_2_5_VERCEL_PRICING: PerTokenPricing = { + const GPT_IMAGE_2_5_VERCEL_AI_GATEWAY_PRICING: PerTokenPricing = { type: "per_token", input: 5, cached_input: 1.25, @@ -1164,10 +1164,10 @@ export namespace models { vercel: { provider: "vercel", id: "openai/gpt-image-2.5-flare", - pricing: GPT_IMAGE_2_5_VERCEL_PRICING, + pricing: GPT_IMAGE_2_5_VERCEL_AI_GATEWAY_PRICING, avg_cost_usd: 0.055, // The model page exposes background=transparent; reference input - // is not established by the gateway feed (2026-09-09 KST). + // is not established by the Vercel AI Gateway feed (2026-09-09 KST). url: "https://vercel.com/ai-gateway/models/gpt-image-2.5-flare", }, openrouter: { @@ -1211,7 +1211,7 @@ export namespace models { options: ["auto", "low", "medium", "high", "xhigh", "max"], default: "high", }, - pricing: GPT_IMAGE_2_5_VERCEL_PRICING, + pricing: GPT_IMAGE_2_5_VERCEL_AI_GATEWAY_PRICING, // High 1024² output estimate is $0.05268, plus a small input allowance. // Not a fixed per-image price: actual cost depends on all billed tokens. // https://developers.openai.com/api/docs/guides/image-generation @@ -1237,10 +1237,10 @@ export namespace models { vercel: { provider: "vercel", id: "openai/gpt-image-2.5-sunburst", - pricing: GPT_IMAGE_2_5_VERCEL_PRICING, + pricing: GPT_IMAGE_2_5_VERCEL_AI_GATEWAY_PRICING, avg_cost_usd: 0.055, // The model page exposes background=transparent; reference input - // is not established by the gateway feed (2026-09-09 KST). + // is not established by the Vercel AI Gateway feed (2026-09-09 KST). url: "https://vercel.com/ai-gateway/models/gpt-image-2.5-sunburst", }, openrouter: { @@ -1282,7 +1282,7 @@ export namespace models { options: ["auto", "low", "medium", "high", "xhigh", "max"], default: "high", }, - pricing: GPT_IMAGE_2_5_VERCEL_PRICING, + pricing: GPT_IMAGE_2_5_VERCEL_AI_GATEWAY_PRICING, avg_cost_usd: 0.055, }, // https://developers.openai.com/api/docs/models/gpt-image-1.5 @@ -1396,7 +1396,7 @@ export namespace models { // Google (multimodal LLMs with native image output) // ----------------------------------------------------------------- // python .tools/model_info.py --image gemini-3.1-flash-image - // Vercel gateway pricing: $0.50/MTok input, $3.00/MTok output + // Vercel AI Gateway pricing: $0.50/MTok input, $3.00/MTok output "google/gemini-3.1-flash-image-preview": { id: "google/gemini-3.1-flash-image-preview", label: "Gemini 3.1 Flash Image", @@ -1410,7 +1410,7 @@ export namespace models { vendor: "google", // "Nano Banana 2"; ids/prices verified 2026-06-29, see issues/908 providers: { - // The gateway serves the graduated `google/gemini-3.1-flash-image` + // Vercel AI Gateway serves the graduated `google/gemini-3.1-flash-image` // and the `-preview` alias at identical rates (feed, 2026-09-02). // Bindings call the graduated id; the canonical key above stays // `-preview` because it is persisted in selections and published @@ -1449,7 +1449,7 @@ export namespace models { avg_cost_usd: 0.004, }, // python .tools/model_info.py --image gemini-3-pro-image - // Vercel gateway pricing: $2.00/MTok input, $12.00/MTok output + // Vercel AI Gateway pricing: $2.00/MTok input, $12.00/MTok output "google/gemini-3-pro-image": { id: "google/gemini-3-pro-image", label: "Gemini 3 Pro Image", @@ -1497,7 +1497,7 @@ export namespace models { }, // "Nano Banana 2 Lite" — GA 2026-06-30. The cost/speed tier of the 3.1 // Flash family: ~half of Nano Banana 2's meter, and 1K-only output - // (2K/4K unsupported — the differentiator). Vercel + OpenRouter both + // (2K/4K unsupported — the differentiator). Vercel AI Gateway + OpenRouter both // meter it at $0.25/$1.50 (verified 2026-07-01); fal id not verified, // so left out. OpenRouter doesn't advertise input_references for the // Lite (t2i only per its model page), so no `references` (TOOL-DESIGN: @@ -1535,7 +1535,7 @@ export namespace models { constraints: { max_edge: 1024 }, pricing: { type: "per_token", input: 0.25, output: 1.5 }, // Published 1K per-image cost (the card default): $0.034 = 1120 tokens - // × $30/1M image-output (Google/Vercel changelog, 2026-07-01). The + // × $30/1M image-output (Google/Vercel AI Gateway changelog, 2026-07-01). The // budget meter charges this per image, so it must be the real cost. avg_cost_usd: 0.034, }, @@ -1555,11 +1555,11 @@ export namespace models { short_description: "Latest Flux model with best-in-class image quality and prompt adherence", vendor: "black-forest-labs", - // All three providers meter $0.03 per megapixel (Vercel model page, + // All three providers meter $0.03 per megapixel (Vercel AI Gateway model page, // OpenRouter endpoint `cost_usd`/megapixel, fal "first megapixel"); - // represented as flat at the 1MP baseline. The Vercel binding shipped + // represented as flat at the 1MP baseline. The Vercel AI Gateway binding shipped // at $0.06 — a 2x hosted over-billing — corrected 2026-09-02. The - // gateway feed carries no `pricing` for BFL cards, so the page is the + // Vercel AI Gateway feed carries no `pricing` for BFL cards, so the page is the // source. providers: { vercel: { @@ -1597,7 +1597,7 @@ export namespace models { // Black Forest Labs — Flux 2 Max // ----------------------------------------------------------------- // BFL's top Flux 2 line (2025-12-16). $0.07 per megapixel on all three - // (Vercel model page — the feed carries no BFL pricing; OpenRouter + // (Vercel AI Gateway model page — the feed carries no BFL pricing; OpenRouter // `cost_usd`/megapixel; fal "first megapixel", +$0.03 each additional). // Represented as flat at the 1MP baseline. Verified 2026-09-02. "bfl/flux-2-max": { @@ -1689,7 +1689,7 @@ export namespace models { }, short_description: "Fast context-aware image generation and editing", vendor: "black-forest-labs", - // $0.04 on both providers: gateway feed `pricing.image` + fal's model + // $0.04 on both providers: Vercel AI Gateway feed `pricing.image` + fal's model // page ("Fixed $0.04 cost per image edit"), verified 2026-09-02. providers: { vercel: { @@ -1744,7 +1744,7 @@ export namespace models { // ----------------------------------------------------------------- // ByteDance — Seedream 5.0 Pro // ----------------------------------------------------------------- - // Vercel feed `pricing.image` $0.035 (the model page's rate table shows + // Vercel AI Gateway feed `pricing.image` $0.035 (the model page's rate table shows // $0.04 while its copy says $0.035 — the feed is the billing contract); // OpenRouter `cost_usd` $0.045 at 1K ($0.09 high-res, +$0.003 per // input image, not modelled); fal $0.0675 for ≤1536² area, $0.135 up @@ -1804,7 +1804,7 @@ export namespace models { // ----------------------------------------------------------------- // ByteDance — Seedream 5.0 Lite // ----------------------------------------------------------------- - // Universal: $0.035/img on all three (Vercel feed `pricing.image`, + // Universal: $0.035/img on all three (Vercel AI Gateway feed `pricing.image`, // OpenRouter `cost_usd`, fal page payload), verified 2026-09-02. A // 2K–4K model: fal scales requests below 2560x1440 up to its floor. "bytedance/seedream-5.0-lite": { @@ -1849,7 +1849,7 @@ export namespace models { styles: null, sizes: null, // fal: total pixels between 2560x1440 and 4096x4096 (requests below - // the floor are scaled up to it); Vercel and OpenRouter serve 2K/4K + // the floor are scaled up to it); Vercel AI Gateway and OpenRouter serve 2K/4K // only. An area envelope, so the default is a 2K request. constraints: { min_pixels: 3_686_400, max_pixels: 16_777_216 }, pricing: { type: "per_image_flat", usd: 0.035 }, @@ -1909,7 +1909,7 @@ export namespace models { // SpaceXAI — Grok Imagine Image 2.0 // ----------------------------------------------------------------- // Tiered by quality (low/medium) × resolution (1K/2K); the same four - // rates on all three providers (Vercel feed + // rates on all three providers (Vercel AI Gateway feed // `image_dimension_quality_pricing`, OpenRouter `cost_usd` variants, // fal page). OpenRouter also bills $0.01 per input image (not // modelled). Verified 2026-09-02. @@ -1992,7 +1992,7 @@ export namespace models { // ----------------------------------------------------------------- // Meta — Muse Image 1.0 // ----------------------------------------------------------------- - // Meta's agentic image model (2026-08-26): $0.01/img on Vercel (feed + // Meta's agentic image model (2026-08-26): $0.01/img on Vercel AI Gateway (feed // `pricing.image`) and fal (page payload). OpenRouter lists it but // exposes no serving endpoint. // fal exposes aspect ratio only (no size control). Verified 2026-09-02. @@ -2036,7 +2036,7 @@ export namespace models { // ----------------------------------------------------------------- // Recraft — V4.1 // ----------------------------------------------------------------- - // Universal: $0.035/img raster on every provider (Vercel feed + // Universal: $0.035/img raster on every provider (Vercel AI Gateway feed // `pricing.image`, OpenRouter endpoint `cost_usd`, fal page payload), // verified 2026-09-02. Vector styles are $0.08 and a separate route on // fal/OpenRouter (`.../text-to-vector`, `recraft-v4.1-vector`) — not @@ -2140,7 +2140,7 @@ export namespace models { * Resolve a model identifier to its cost card (data only). * * Accepts: - * - Full gateway id (`"bfl/flux-pro-1.1"`) + * - Full provider model id (`"bfl/flux-pro-1.1"`) * - The deprecated `ProviderModel` wrapper * - Bare provider id (`"flux-pro-1.1"`) — exact match against the * segment after the `vendor/` prefix. Unlike @@ -2940,7 +2940,7 @@ export namespace models { /** * A provider that can serve a video model. Distinct from the top-level * {@link models.Provider} because video routes through more than the - * Vercel gateway. Each provider uses its own id format and meter. + * Vercel AI Gateway. Each provider uses its own id format and meter. */ export type VideoProvider = (typeof providers)[number]; @@ -2965,7 +2965,7 @@ export namespace models { /** * Whether a clip is generated with synchronized audio. A real pricing axis: - * both fal and the Vercel gateway meter `silent` at roughly half of + * both fal and the Vercel AI Gateway meter `silent` at roughly half of * `audio`; Seedance bundles audio into its single rate. */ export type AudioMode = "audio" | "silent"; @@ -2979,7 +2979,7 @@ export namespace models { * A binding lists only the `(resolution, mode)` combinations its provider * actually serves and meters, so the keys double as that provider's * resolution/audio support: Seedance lists only `audio` because it bundles - * audio into one rate, and Veo 3.1 Lite omits `"4k"` because the gateway + * audio into one rate, and Veo 3.1 Lite omits `"4k"` because Vercel AI Gateway * does not sell it at that line. Each value is the real * USD-per-output-second rate for that exact config. * @@ -3018,7 +3018,7 @@ export namespace models { provider: VideoProvider; /** * Provider-specific call id. Format varies — - * `google/veo-3.1-generate-001` (Vercel), `fal-ai/veo3.1/image-to-video` + * `google/veo-3.1-generate-001` (Vercel AI Gateway), `fal-ai/veo3.1/image-to-video` * (fal, where the capability is keyed into the endpoint id). */ id: string; @@ -3097,9 +3097,9 @@ export namespace models { speed_label: "slow", url: "https://deepmind.google/models/veo/", providers: { - // Vercel AI Gateway — gateway.video(id), image-to-video. The - // gateway meters both audio modes and sells 4K; the matrix is - // identical to fal's. Verified against the gateway's own + // Vercel AI Gateway — vercelAiGateway.videoModel(id), image-to-video. The + // Vercel AI Gateway meters both audio modes and sells 4K; the matrix is + // identical to fal's. Verified against Vercel AI Gateway's own // /v1/models feed (`video_duration_pricing`) on 2026-09-02 — the // previous card claimed "audio-on only, ≤1080p", which the feed // contradicts. @@ -3128,7 +3128,7 @@ export namespace models { }, // fal.ai — image-to-video endpoint (capability is keyed into the id; // t2v is a separate `fal-ai/veo3.1` endpoint, not catalogued). Rate - // matrix is identical to the Vercel gateway's: $0.40/s audio and + // matrix is identical to the Vercel AI Gateway's: $0.40/s audio and // $0.20/s silent at 720p/1080p, $0.60/$0.40 at 4K. // https://fal.ai/models/fal-ai/veo3.1/image-to-video fal: { @@ -3169,8 +3169,8 @@ export namespace models { // Google — Veo 3.1 Fast // ----------------------------------------------------------------- // Same envelope as Veo 3.1 (16:9/9:16, 4/6/8s, native audio, ≤4K) at - // ~2.7x lower cost. Rate matrix is identical on Vercel and fal — - // gateway feed `video_duration_pricing` and fal's stated per-second + // ~2.7x lower cost. Rate matrix is identical on Vercel AI Gateway and fal — + // Vercel AI Gateway feed `video_duration_pricing` and fal's stated per-second // rates, verified 2026-09-02. "google/veo-3.1-fast": { id: "google/veo-3.1-fast", @@ -3228,8 +3228,8 @@ export namespace models { // Google — Veo 3.1 Lite // ----------------------------------------------------------------- // The budget Veo: 720p/1080p only, 4/6/8s, native audio. Rates - // identical on Vercel and fal (verified 2026-09-02). Vercel supports - // both text and image input (FAQ verified 2026-09-07); the fal binding + // identical on Vercel AI Gateway and fal (verified 2026-09-02). Vercel AI Gateway + // supports both text and image input (FAQ verified 2026-09-07); the fal binding // below requires an image. Hosted GG's narrower wire is a host concern. "google/veo-3.1-lite": { id: "google/veo-3.1-lite", @@ -3285,9 +3285,9 @@ export namespace models { // ----------------------------------------------------------------- // Alibaba — Wan 3.0 // ----------------------------------------------------------------- - // Per-second by resolution, audio bundled into the rate (the gateway + // Per-second by resolution, audio bundled into the rate (Vercel AI Gateway // feed has no audio axis; fal exposes an `audio` toggle but bills the - // same). Identical on Vercel and fal, verified 2026-09-02. 2–30s. + // same). Identical on Vercel AI Gateway and fal, verified 2026-09-02. 2–30s. "alibaba/wan-3.0": { id: "alibaba/wan-3.0", label: "Wan 3.0", @@ -3361,14 +3361,14 @@ export namespace models { audio: true, speed_label: "slow", url: "https://seed.bytedance.com/en/seedance2_0", - // NO Vercel binding, deliberately. The gateway serves + // NO Vercel AI Gateway binding, deliberately. Vercel AI Gateway serves // `bytedance/seedance-2.0` but meters it PER TOKEN // (`video_token_pricing`: $7.00/MTok at 480p/720p, $7.70 at 1080p, // $4.00 at 4K; reduced with video input; "minimum token floors based // on output duration"). `PerSecondPricing` cannot express that, the // hosted route pre-prices rate×duration, and ByteDance publishes no // tokens-per-second figure — so there is no honest per-second rate, - // and the per-second Vercel binding this card shipped with was + // and the per-second Vercel AI Gateway binding this card shipped with was // invented. Withheld until the hosted path can meter post-flight; // hosted requests fail with "not available on the hosted provider". // See gridaco/grida#1019 (Blocker A). Feed checked 2026-09-02. @@ -3394,7 +3394,7 @@ export namespace models { }, // OpenRouter — async `/api/v1/videos`. Flat $0.06726/s (verified // 2026-06-29, https://openrouter.ai/bytedance/seedance-2.0) — far - // below Vercel's per-resolution rate (the proprietary-pricing- + // below Vercel AI Gateway's per-resolution rate (the proprietary-pricing- // diverges finding, #908). No separate silent meter surfaced. openrouter: { provider: "openrouter", @@ -3416,11 +3416,11 @@ export namespace models { // ByteDance — Seedance 2.5 // ----------------------------------------------------------------- // Newer generation (2026-07-31), NOT a drop-in successor: ~55% more - // per token on the gateway ($10.70/MTok at 480p/720p, $11.70 at 1080p + // per token on Vercel AI Gateway ($10.70/MTok at 480p/720p, $11.70 at 1080p // vs 2.0's $7.00/$7.70), ~56–71% more per second on fal, and no 4K. // It buys 4–30s clips, video-editing and extend-video. 2.0 is cheaper - // and serves 4K. The gateway bills tokens, so there is no comparable - // per-second Vercel meter here. fal states per-second prices, audio bundled. + // and serves 4K. Vercel AI Gateway bills tokens, so there is no comparable + // per-second Vercel AI Gateway meter here. fal states per-second prices, audio bundled. "bytedance/seedance-2.5": { id: "bytedance/seedance-2.5", label: "Seedance 2.5", @@ -3464,7 +3464,7 @@ export namespace models { // SpaceXAI — Grok Imagine Video 1.5 // ----------------------------------------------------------------- // Image-to-video only (no t2v, per SpaceXAI docs); native lip-synced audio - // bundled into the rate. Per-second by resolution, identical on Vercel + // bundled into the rate. Per-second by resolution, identical on Vercel AI Gateway // (no markup) and fal: $0.08/s @480p, $0.14/s @720p, $0.25/s @1080p. // Both also bill $0.01 per input image, captured separately from the // output meter. @@ -3487,7 +3487,7 @@ export namespace models { url: "https://docs.x.ai/developers/models/grok-imagine-video-1.5", providers: { // Vercel AI Gateway — image-to-video; mirrors SpaceXAI's list price (no markup). - // The gateway namespaces every SpaceXAI model under `spacexai/`, so the + // Vercel AI Gateway namespaces every SpaceXAI model under `spacexai/`, so the // call id deliberately differs from this card's canonical `xai/` id. // https://vercel.com/changelog/grok-imagine-video-1-5-on-ai-gateway vercel: { diff --git a/packages/grida-ai/README.md b/packages/grida-ai/README.md index aed089db61..1dcc55ae19 100644 --- a/packages/grida-ai/README.md +++ b/packages/grida-ai/README.md @@ -1,7 +1,7 @@ # @grida/ai Private, experimental SDK for shared model-driven operations. Catalogue-backed -image and video generation use BYOK OpenRouter, Vercel, fal, or scoped Grida +image and video generation use BYOK OpenRouter, Vercel AI Gateway, fal, or scoped Grida Gateway credentials. Music generation uses the existing GG-only Lyria route; sound effects and speech use their existing ElevenLabs BYOK operations. Three existing fal 3D endpoints and direct Tripo H3.1/P1/P2 model generation return a @@ -112,7 +112,7 @@ Objects reject extra fields. Optional fields may be omitted; explicit `null` and is capped at 16 before generation can acquire authority or submit batches. Text normalization remains operation-specific: music uses trimmed UTF-16 length; SFX and 3D use trimmed code-point length; image, video, and speech preserve input -text. Direct Vercel video rejects `seed: 0` because its pinned upstream serializer +text. Direct Vercel AI Gateway video rejects `seed: 0` because its pinned upstream serializer drops zero. The inspected schema describes accepted SDK fields, not every option or value advertised by a provider's model card. @@ -161,7 +161,7 @@ const { images: generated } = await operation.generate({ ``` `resolve` requires both a catalogue model ID and a provider. `provider: "auto"` -explicitly adopts the existing order: connected OpenRouter, Vercel, fal, then GG. +explicitly adopts the existing order: connected OpenRouter, Vercel AI Gateway, fal, then GG. Explicit choices check only that provider. The operation freezes the selected model, binding, provider, and reference cap. Generation reads the selected provider's key again; losing that credential fails without switching providers. @@ -187,7 +187,7 @@ numeric `aspect_ratio`, integer `seed`, optional `quality`, and `signal`, subjec to the selected endpoint's input schema. GPT Image 2.5's OpenRouter and fal routes do not accept seed; fal also rejects aspect ratio and uses explicit dimensions or auto size. fal edits use the distinct catalogue edit binding with 1–16 references. -Explicit quality values, including `auto`, are forwarded; Vercel's OpenAI models +Explicit quality values, including `auto`, are forwarded; Vercel AI Gateway's OpenAI models use the `openai` namespace. Provider batch limits remain authoritative: a requested count may require multiple submissions. Every batch sets **`maxRetries: 0`**. Failed generation is never automatically resubmitted; fal status polling is separate. @@ -230,8 +230,8 @@ where that exact operation's input schema declares `image`. A text selection rej image inputs. Capability facts are owned by `@grida/ai-models`; absent facts in old snapshots use only exact canonical/provider/binding matches to bundled facts. Changed, removed, or explicitly unknown bindings do not inherit capabilities. -GG remains text-only and requires a text-eligible Vercel binding. The current fal -bindings and Vercel Grok require a start frame. OpenRouter uses `frame_images` with +GG remains text-only and requires a text-eligible Vercel AI Gateway binding. The current fal +bindings and Vercel AI Gateway Grok require a start frame. OpenRouter uses `frame_images` with `first_frame`; the exact fal Wan binding uses `start_image_url`. The exact fal `fal-ai/veo3.1/lite/image-to-video` binding (bundled as @@ -250,7 +250,7 @@ The [fal API documents inline file inputs and an 8 MB image limit](https://fal.a and the [model input form lists the admitted formats](https://fal.ai/models/fal-ai/veo3.1/lite/image-to-video). The SDK uses this conservative format subset and decimal byte ceiling; it does not decode image pixels, resize frames, or promise upstream content acceptance. -Other video bindings remain HTTPS-only, including OpenRouter and Vercel. +Other video bindings remain HTTPS-only, including OpenRouter and Vercel AI Gateway. `MediaOperations.inspect(...).input_schema.properties.image` is present only when bytes are supported. Its nested `data` schema states the base64 and decoded @@ -269,15 +269,15 @@ the same native/JSON parser that owns serialization eligibility. Other routes accept `aspect_ratio` as positive integer `W:H`, `resolution` as positive integer `WxH` (not a catalogue price label such as `720p`), positive finite -`duration`/`fps`, and a safe-integer `seed`. The pinned Vercel adapter silently omits -zero upstream, so direct Vercel `seed: 0` is rejected as `invalid_input` before key +`duration`/`fps`, and a safe-integer `seed`. The pinned Vercel AI Gateway adapter silently omits +zero upstream, so direct Vercel AI Gateway `seed: 0` is rejected as `invalid_input` before key lookup or submission; the other adapters preserve zero. It requests **one video once**. There is no arbitrary provider-options object, raw SDK model, or retry setting. Queue status reads do not resubmit the job. Multiple returned videos remain supported within a 16-item, 64 MiB decoded aggregate bound. Submission/poll JSON, GG/SSE encoded envelopes, inline data, authenticated OpenRouter content, and result -download streams are bounded before retention or decoding. Vercel result URLs -must use its exact gateway origin or inline data; fal queue/result URLs stay on +download streams are bounded before retention or decoding. Vercel AI Gateway result URLs +must use its exact Vercel AI Gateway origin or inline data; fal queue/result URLs stay on its allowed hosts. OpenRouter's authenticated content endpoint is used instead of provider-advertised unsigned URLs. Hosts still authorize DNS and redirect hops. diff --git a/packages/grida-ai/src/image-byok.test.ts b/packages/grida-ai/src/image-byok.test.ts index c8315ef6ef..463e817bc4 100644 --- a/packages/grida-ai/src/image-byok.test.ts +++ b/packages/grida-ai/src/image-byok.test.ts @@ -681,11 +681,11 @@ describe("makeImageModelFor", () => { expect(m.provider).toBe("openrouter"); }); - it("builds a vercel image model without throwing", () => { + it("builds a Vercel AI Gateway image model without throwing", () => { expect(makeImageModelFor("vercel", "sk", "bfl/flux-2-pro")).toBeTruthy(); }); - it("keeps Vercel Gateway image results on the provider request lane", async () => { + it("keeps Vercel AI Gateway image results on the provider request lane", async () => { const request = vi.fn< (input: string | URL | Request, init?: MockInit) => Promise >(async (input: string | URL | Request, init: MockInit = {}) => { @@ -702,7 +702,9 @@ describe("makeImageModelFor", () => { ); }); const download = vi.fn(async () => { - throw new Error("Vercel image results must never open a download lane"); + throw new Error( + "Vercel AI Gateway image results must never open a download lane" + ); }); const model = makeImageModelFor( "vercel", diff --git a/packages/grida-ai/src/image-byok.ts b/packages/grida-ai/src/image-byok.ts index 7ed20e1a45..361c337020 100644 --- a/packages/grida-ai/src/image-byok.ts +++ b/packages/grida-ai/src/image-byok.ts @@ -1,7 +1,7 @@ // GRIDA-SEC-004 — provider-owned submissions and credential-free result downloads. /** Internal AI SDK adapters; ImageClient owns the safe public operation boundary. */ -import { createGateway } from "@ai-sdk/gateway"; +import { createGateway as createVercelAiGateway } from "@ai-sdk/gateway"; import type { ImageModelV3, ImageModelV3CallOptions } from "@ai-sdk/provider"; import type { models } from "@grida/ai-models"; import type { ImageClient } from "./image-client"; @@ -24,12 +24,15 @@ function makeOpenRouterImageModel( return new OpenRouterImageModel(apiKey, id, providerHttp); } -function makeVercelImageModel( +function makeVercelAiGatewayImageModel( apiKey: string, id: string, providerHttp: ProviderHttp ): ImageModelV3 { - return createGateway({ apiKey, fetch: providerHttp.request }).imageModel(id); + return createVercelAiGateway({ + apiKey, + fetch: providerHttp.request, + }).imageModel(id); } function makeFalImageModel( @@ -57,7 +60,7 @@ export function makeImageModelFor( model = makeOpenRouterImageModel(apiKey, id, providerHttp); break; case "vercel": - model = makeVercelImageModel(apiKey, id, providerHttp); + model = makeVercelAiGatewayImageModel(apiKey, id, providerHttp); break; case "fal": model = makeFalImageModel(apiKey, id, providerHttp); @@ -94,8 +97,8 @@ class BackgroundImageModel implements ImageModelV3 { } doGenerate(options: ImageModelV3CallOptions) { - // Gateway forwards providerOptions unchanged; native OpenAI settings live - // under openai, not vercel. Catalogue admission is checked independently. + // Vercel AI Gateway forwards providerOptions unchanged; native OpenAI settings live + // under the native openai namespace. Catalogue admission is checked independently. const namespace = this.routeProvider === "vercel" ? "openai" : this.routeProvider; return this.model.doGenerate({ diff --git a/packages/grida-ai/src/image-client.test.ts b/packages/grida-ai/src/image-client.test.ts index 3b10c94aa3..8d1659f641 100644 --- a/packages/grida-ai/src/image-client.test.ts +++ b/packages/grida-ai/src/image-client.test.ts @@ -103,7 +103,7 @@ describe("ImageClient public operations", () => { ); }); - it("uses Vercel's declared image protocol with the supplied key", async () => { + it("uses Vercel AI Gateway's declared image protocol with the supplied key", async () => { const { client, request, download } = setup(); request.mockImplementation(async () => Response.json({ images: [BASE64] })); const operation = await client.resolve({ diff --git a/packages/grida-ai/src/media-operations.test.ts b/packages/grida-ai/src/media-operations.test.ts index 8a6985a87b..ec16c59025 100644 --- a/packages/grida-ai/src/media-operations.test.ts +++ b/packages/grida-ai/src/media-operations.test.ts @@ -435,7 +435,7 @@ describe("MediaOperations JSON input", () => { ); }); - it("requires the video frame only in its selected variant and preserves the Vercel zero-seed restriction", () => { + it("requires the video frame only in its selected variant and preserves the Vercel AI Gateway zero-seed restriction", () => { expect(properties(operations.inspect(video)).seed).toMatchObject({ not: { const: 0 }, }); diff --git a/packages/grida-ai/src/provider-credentials.test.ts b/packages/grida-ai/src/provider-credentials.test.ts index fa5cfb595d..ddbd2e984f 100644 --- a/packages/grida-ai/src/provider-credentials.test.ts +++ b/packages/grida-ai/src/provider-credentials.test.ts @@ -126,7 +126,7 @@ describe("ProviderCredentials.normalize", () => { } ); - it("admits current and opaque legacy Vercel keys without declaring them authenticated", () => { + it("admits current and opaque legacy Vercel AI Gateway keys without declaring them authenticated", () => { for (const value of [ "vck_x", "vck_nonstandard!suffix", @@ -216,7 +216,7 @@ describe("ProviderCredentials.check", () => { expect(download).not.toHaveBeenCalled(); }); - it("accepts a Vercel legacy key only when the explicit check succeeds", async () => { + it("accepts a Vercel AI Gateway legacy key only when the explicit check succeeds", async () => { const { client } = setup(async () => Response.json(successes.vercel)); await expect( client.check({ provider: "vercel", key: "opaque-legacy-key" }) diff --git a/packages/grida-ai/src/provider-ids.ts b/packages/grida-ai/src/provider-ids.ts index 837603aa4a..3a59e19895 100644 --- a/packages/grida-ai/src/provider-ids.ts +++ b/packages/grida-ai/src/provider-ids.ts @@ -1,7 +1,7 @@ // GRIDA-GG: provider — shared provider identities and explicit legacy precedence. /** * Which generation modalities a BYOK provider serves. A provider may serve - * several: OpenRouter does text + image; Vercel does text + image + video; fal + * several: OpenRouter does text + image; Vercel AI Gateway does text + image + video; fal * does image + video. This vocabulary is intentionally limited to the generic * model resolvers that consume {@link byokProvidersFor}. Provider-shaped * endpoints such as fal 3D generation and ElevenLabs Sound Effects/Text to @@ -18,7 +18,7 @@ export const BYOK_PROVIDER_METADATA = [ }, { id: "vercel", - label: "Vercel", + label: "Vercel AI Gateway", modalities: ["text", "image", "video"], }, { diff --git a/packages/grida-ai/src/video-client.test.ts b/packages/grida-ai/src/video-client.test.ts index 08022f1bc0..f1821e9eca 100644 --- a/packages/grida-ai/src/video-client.test.ts +++ b/packages/grida-ai/src/video-client.test.ts @@ -427,7 +427,7 @@ describe("VideoClient bounded image input", () => { }); describe("VideoClient public operation", () => { - it("freezes explicit selection without a credential/model access surface and uses Vercel's video wire", async () => { + it("freezes explicit selection without a credential/model access surface and uses Vercel AI Gateway's video wire", async () => { const { client, request, get, download } = setup(); const operation = await client.resolve({ model_id: ID, @@ -728,7 +728,7 @@ describe("VideoClient public operation", () => { "https://other.example/result.mp4", "https://ai-gateway.vercel.sh.evil.example/result.mp4", "https://ai-gateway.vercel.sh:8443/result.mp4", - ])("refuses an untrusted Vercel result origin %s", async (url) => { + ])("refuses an untrusted Vercel AI Gateway result origin %s", async (url) => { const { client, request, download } = setup(); request.mockImplementation(async () => sse([{ type: "url", url, mediaType: "video/mp4" }]) @@ -747,7 +747,7 @@ describe("VideoClient public operation", () => { it.each([ "https://ai-gateway.vercel.sh/result.mp4", "data:video/mp4;base64,AQID", - ])("accepts allowed Vercel results %s", async (url) => { + ])("accepts allowed Vercel AI Gateway results %s", async (url) => { const { client, request, download } = setup(); request.mockImplementation(async () => sse([{ type: "url", url, mediaType: "video/mp4" }]) @@ -843,7 +843,7 @@ describe("VideoClient public operation", () => { expect(request).toHaveBeenCalledOnce(); }); - it("refuses the pinned Vercel adapter's unrepresentable zero seed before key lookup", async () => { + it("refuses the pinned Vercel AI Gateway adapter's unrepresentable zero seed before key lookup", async () => { const { client, request, get } = setup(); const operation = await client.resolve({ model_id: ID, @@ -1051,7 +1051,7 @@ describe("video execution and result bounds", () => { expect(response.body!.locked).toBe(false); }); - it("cancels a Vercel SSE connection after its first result even when the server leaves it open", async () => { + it("cancels a Vercel AI Gateway SSE connection after its first result even when the server leaves it open", async () => { const { client, request } = setup(); const cancel = vi.fn<() => void>(); request.mockImplementation( diff --git a/packages/grida-ai/src/video-client.ts b/packages/grida-ai/src/video-client.ts index ed27b5e467..8a7c4eee4c 100644 --- a/packages/grida-ai/src/video-client.ts +++ b/packages/grida-ai/src/video-client.ts @@ -246,7 +246,7 @@ export namespace VideoClient { fps?: number; /** Explicit audio control only where the selected route's schema advertises it. */ generate_audio?: boolean; - /** Safe integer. The pinned Vercel adapter cannot honor zero and rejects it. */ + /** Safe integer. The pinned Vercel AI Gateway adapter cannot honor zero and rejects it. */ seed?: number; /** An already authorized HTTPS start frame. Mutually exclusive with image. */ image_url?: string; @@ -333,7 +333,7 @@ function resultUrl(value: string, provider: VideoClient.Provider): URL { throw new VideoClient.Failure("invalid_response"); if ( provider === "vercel" && - url.origin !== new URL(videoModels.vercelBase).origin + url.origin !== new URL(videoModels.vercelAiGatewayVideoBaseUrl).origin ) throw new VideoClient.Failure("unsupported_untrusted_result_origin"); return url; diff --git a/packages/grida-ai/src/video-models.ts b/packages/grida-ai/src/video-models.ts index 3ee8065a9d..fbdaa1f39a 100644 --- a/packages/grida-ai/src/video-models.ts +++ b/packages/grida-ai/src/video-models.ts @@ -1,6 +1,6 @@ // GRIDA-SEC-004 / GRIDA-SEC-006 — fixed provider video wires and scoped GG submission. // GRIDA-GG: token — hosted video is text-only and uses the shared live-token contract. -import { createGateway } from "@ai-sdk/gateway"; +import { createGateway as createVercelAiGateway } from "@ai-sdk/gateway"; import { assertAllowedUrl, falQueueOutcome, pollQueue } from "./fetch-helpers"; import { postHosted } from "./gg"; import type { GgTokenSource } from "./gg-session"; @@ -12,7 +12,8 @@ import type { VideoClient } from "./video-client"; export namespace videoModels { export const maxBytes = MediaInputs.limits.video; export const maxEnvelopeBytes = Math.ceil(maxBytes / 3) * 4 + 64 * 1024; - export const vercelBase = "https://ai-gateway.vercel.sh/v3/ai"; + export const vercelAiGatewayVideoBaseUrl = + "https://ai-gateway.vercel.sh/v3/ai"; const falHosts = ["fal.run", "*.fal.run", "fal.media", "*.fal.media"]; const openRouterBase = "https://openrouter.ai/api/v1/videos"; export type Input = { @@ -41,9 +42,9 @@ export namespace videoModels { request.check(); if (provider === "fal") return fal(key, id, input, request); if (provider === "openrouter") return openRouter(key, id, input, request); - const model = createGateway({ + const model = createVercelAiGateway({ apiKey: key, - baseURL: vercelBase, + baseURL: vercelAiGatewayVideoBaseUrl, fetch: request.transport(videoModels.maxEnvelopeBytes).request, }).videoModel(id); // Call the model directly: the high-level SDK adds retries and automatic URL downloads. diff --git a/packages/grida-cli/README.md b/packages/grida-cli/README.md index e8c7635d28..6a2265d580 100644 --- a/packages/grida-cli/README.md +++ b/packages/grida-cli/README.md @@ -120,7 +120,7 @@ without Grida login. `providers configure ` stores a key using hidden or `--key-stdin`; `providers remove ` removes the shared stored key. All selected keys (file, environment or stdin) pass cheap static validation through [the shared provider policy](https://github.com/gridaco/grida/blob/main/packages/grida-ai/README.md). -`configure` additionally checks OpenRouter, Vercel, fal and Tripo once before saving; +`configure` additionally checks OpenRouter, Vercel AI Gateway, fal and Tripo once before saving; rejection, denial or an unavailable check leaves the old key unchanged. ElevenLabs has no suitable permission-neutral check and saves with `verification.status` set to `not_supported`. Successful supported checks report `accepted`, which proves diff --git a/packages/grida-cli/src/cli.ts b/packages/grida-cli/src/cli.ts index 23bf0f6672..986e1e0223 100644 --- a/packages/grida-cli/src/cli.ts +++ b/packages/grida-cli/src/cli.ts @@ -34,7 +34,7 @@ export namespace Cli { providers: "Usage: grida providers \n\nCommands:\n list Show key presence and effective source\n configure Save a shared provider API key\n remove Remove a shared provider API key", "providers configure": - "Usage: grida providers configure [--key-stdin] [--json] [--no-input]\n\nSave an API key in shared plaintext credentials.toml with private permissions.\nDesktop and CLI use the same stored keys. No Grida login required.\nValidate format, then check OpenRouter/Vercel/fal/Tripo once before saving.\nA failed check leaves stored keys unchanged. ElevenLabs saves unverified.\nWithout --key-stdin, enter the key at a hidden terminal prompt.\nAutomation requires --key-stdin; keys are never accepted as arguments.", + "Usage: grida providers configure [--key-stdin] [--json] [--no-input]\n\nSave an API key in shared plaintext credentials.toml with private permissions.\nDesktop and CLI use the same stored keys. No Grida login required.\nValidate format, then check OpenRouter/Vercel AI Gateway/fal/Tripo once before saving.\nA failed check leaves stored keys unchanged. ElevenLabs saves unverified.\nWithout --key-stdin, enter the key at a hidden terminal prompt.\nAutomation requires --key-stdin; keys are never accepted as arguments.", "providers remove": "Usage: grida providers remove [--json] [--no-input]\n\nRemove the stored key for Desktop and CLI. Environment keys remain effective.\nThis does not revoke the key at its provider or sign out of Grida.", "providers list": diff --git a/packages/grida-cli/src/media-http.test.ts b/packages/grida-cli/src/media-http.test.ts index a2ccb4b6f4..ca42b6fe8b 100644 --- a/packages/grida-cli/src/media-http.test.ts +++ b/packages/grida-cli/src/media-http.test.ts @@ -259,7 +259,7 @@ describe("MediaHttp provider authority", () => { "https://ai-gateway.vercel.sh/v3/ai/image-model", "POST", { - authorization: "Bearer synthetic-gateway", + authorization: "Bearer synthetic-vercel-ai-gateway", "content-type": "application/json", "ai-model-id": "fixture", "ai-image-model-specification-version": "3", @@ -272,7 +272,7 @@ describe("MediaHttp provider authority", () => { "https://ai-gateway.vercel.sh/v3/ai/video-model", "POST", { - authorization: "Bearer synthetic-gateway", + authorization: "Bearer synthetic-vercel-ai-gateway", "content-type": "application/json", "ai-video-model-specification-version": "3", }, diff --git a/packages/grida-cli/src/media-http.ts b/packages/grida-cli/src/media-http.ts index 5270ddd069..a48e1c53d5 100644 --- a/packages/grida-cli/src/media-http.ts +++ b/packages/grida-cli/src/media-http.ts @@ -241,13 +241,13 @@ function snapshot( : destination !== "download" && destination !== "tripo-upload" && name === "authorization"; - const gateway = + const isVercelAiGatewayHeader = destination === "vercel" && (name === "ai-gateway-protocol-version" || name === "ai-gateway-auth-method" || name === "ai-model-id" || /^ai-(?:image|video)-model-specification-version$/.test(name)); - if (!common && !credential && !gateway) fail(); + if (!common && !credential && !isVercelAiGatewayHeader) fail(); } if (destination === "download") { if ( diff --git a/packages/grida-cli/src/provider-credentials.test.ts b/packages/grida-cli/src/provider-credentials.test.ts index fd0de03455..aedec60159 100644 --- a/packages/grida-cli/src/provider-credentials.test.ts +++ b/packages/grida-cli/src/provider-credentials.test.ts @@ -24,7 +24,7 @@ describe("ProviderCredentials environment", () => { it("reads only the exact supported environment names", async () => { const env = { OPENROUTER_API_KEY: " sk-or-openrouter-synthetic \n", - AI_GATEWAY_API_KEY: "vck_gateway-synthetic", + AI_GATEWAY_API_KEY: "vck_vercel-ai-gateway-synthetic", FAL_KEY: "fal:synthetic", ELEVENLABS_API_KEY: "elevenlabs-synthetic", TRIPO_API_KEY: "tsk_synthetic-tripo", @@ -72,7 +72,7 @@ describe("ProviderCredentials environment", () => { }, ]); expect(owner.get("openrouter")).toBe("sk-or-openrouter-synthetic"); - expect(owner.get("vercel")).toBe("vck_gateway-synthetic"); + expect(owner.get("vercel")).toBe("vck_vercel-ai-gateway-synthetic"); expect(owner.get("fal")).toBe("fal:synthetic"); expect(owner.get("elevenlabs")).toBe("elevenlabs-synthetic"); expect(owner.get("tripo")).toBe("tsk_synthetic-tripo"); diff --git a/scripts/ai-local/README.md b/scripts/ai-local/README.md index 973c35ff14..0745d4f55b 100644 --- a/scripts/ai-local/README.md +++ b/scripts/ai-local/README.md @@ -34,7 +34,7 @@ agent, daemon, auth, account, CLI, Desktop and framework packages are unavailabl The factual `@grida/ai-models` root and service `@grida/ai-models/grida` entry must both load in ESM and CommonJS. Snapshot codecs and types come from the service entry; the SDK retains its bundled service defaults. -Synthetic transports exercise OpenRouter, Vercel, fal and scoped GG image +Synthetic transports exercise OpenRouter, Vercel AI Gateway, fal and scoped GG image generation and video submit/poll/result chains, input capability discovery, credential-free result downloads, explicit provider selection, missing credentials, single submission on failure, safe error projection, cleared GG authority, and video cancellation during a diff --git a/scripts/cli-media-local/README.md b/scripts/cli-media-local/README.md index d8e3f3ad92..f418fc6116 100644 --- a/scripts/cli-media-local/README.md +++ b/scripts/cli-media-local/README.md @@ -59,7 +59,7 @@ cases verify file admission and wire serialization; provider codec acceptance remains a separate check. Provider configuration uses distinct synthetic keys and the real SDK credential -policy. OpenRouter, Vercel and fal each make exactly one authenticated registration +policy. OpenRouter, Vercel AI Gateway and fal each make exactly one authenticated registration read against their fixed synthetic key/credits/pricing response. ElevenLabs makes no registration request and reports `not_supported`. Accepted checks report only safe verification metadata; they do not establish remaining credits or future diff --git a/scripts/cli-media-local/proof.mjs b/scripts/cli-media-local/proof.mjs index 3013c8dea5..b85e1b190a 100644 --- a/scripts/cli-media-local/proof.mjs +++ b/scripts/cli-media-local/proof.mjs @@ -32,7 +32,7 @@ const prompt = "A synthetic media fixture"; const keys = { openrouter: "sk-or-synthetic-openrouter-key-no-authority", vercel: - "vck_syntheticVercelKeyNoAuthority0123456789abcdefghijklmnopqrstuvwxyz", + "vck_syntheticVercelAiGatewayKeyNoAuthority0123456789abcdefghijklmnopqrstuvwxyz", fal: "synthetic-fal-id:synthetic-fal-secret-no-authority", elevenlabs: "synthetic-elevenlabs-key-no-authority", }; @@ -1015,34 +1015,37 @@ async function main() { ); } ); - await check("Vercel image uses the pinned SDK protocol", async () => { - await generate({ - name: "vercel-image", - model: "openai/gpt-image-2", - provider: "vercel", - value: { prompt }, - extraEnv: { AI_GATEWAY_API_KEY: keys.vercel }, - wire: fixture([ - { - hostname: "ai-gateway.vercel.sh", - path: "/v3/ai/image-model", - method: "POST", - headers: { - authorization: `Bearer ${keys.vercel}`, - "ai-model-id": "openai/gpt-image-2", - "ai-image-model-specification-version": "3", + await check( + "Vercel AI Gateway image uses the pinned SDK protocol", + async () => { + await generate({ + name: "vercel-ai-gateway-image", + model: "openai/gpt-image-2", + provider: "vercel", + value: { prompt }, + extraEnv: { AI_GATEWAY_API_KEY: keys.vercel }, + wire: fixture([ + { + hostname: "ai-gateway.vercel.sh", + path: "/v3/ai/image-model", + method: "POST", + headers: { + authorization: `Bearer ${keys.vercel}`, + "ai-model-id": "openai/gpt-image-2", + "ai-image-model-specification-version": "3", + }, + json: { prompt, n: 1, providerOptions: {} }, + response: jsonBody({ + images: [png.toString("base64")], + warnings: [], + }), }, - json: { prompt, n: 1, providerOptions: {} }, - response: jsonBody({ - images: [png.toString("base64")], - warnings: [], - }), - }, - ]), - data: png, - type: "image/png", - }); - }); + ]), + data: png, + type: "image/png", + }); + } + ); await check("ElevenLabs SFX accepts explicit stdin key", async () => { await generate({ name: "sfx", diff --git a/test/billing-quota-and-ai.md b/test/billing-quota-and-ai.md index 1fc8c64e59..44d054e1b6 100644 --- a/test/billing-quota-and-ai.md +++ b/test/billing-quota-and-ai.md @@ -212,7 +212,7 @@ A contributor sets a `BYOK_*` key (e.g. `BYOK_OPENROUTER_API_KEY`). AI call goes ### TC-BILLING-AI-036 — BYOK key accidentally set on a hosted deploy Bug: a `BYOK_*` secret leaks into a hosted/preview deploy. -**Expected:** There is **no code-level production guard** (by design — same server-env trust model as `OPENAI_API_KEY` / `REPLICATE_API_TOKEN`). With the key set, every org's calls bypass billing and the org-id sanity gate (auth still holds). This is a documented residual risk (see `SECURITY.md` GRIDA-SEC-003), mitigated operationally by never setting the secret in the hosted product — not by code. There is no "production check" to fail. +**Expected:** There is **no code-level production guard** (by design — same trust model as other server-only credentials). With the key set, every org's calls bypass billing and the org-id sanity gate (auth still holds). This is a documented residual risk (see `SECURITY.md` GRIDA-SEC-003), mitigated operationally by never setting the secret in the hosted product — not by code. There is no "production check" to fail. ### TC-BILLING-AI-037 — Disabled model attempted diff --git a/turbo.json b/turbo.json index 550f8adaf6..9d91923705 100644 --- a/turbo.json +++ b/turbo.json @@ -7,13 +7,14 @@ "BIRD_WORKSPACE_ID", "EDGE_CONFIG", "ENABLE_EXPERIMENTAL_COREPACK", + "GG_REPLICATE_API_TOKEN", + "GG_VERCEL_AI_GATEWAY_API_KEY", "GRIDA_S2S_PRIVATE_API_KEY", "INTEGRATIONS_TEST_TOSSPAYMENTS_CUSTOMER_KEY", "INTEGRATIONS_TEST_TOSSPAYMENTS_SECRET_KEY", "IPINFO_ACCESS_TOKEN", "OPENAI_API_KEY", "REDIS_URL", - "REPLICATE_API_TOKEN", "RESEND_API_KEY", "SENTRY_AUTH_TOKEN", "SENTRY_HOST",