Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 20 additions & 6 deletions SECURITY.md
Original file line number Diff line number Diff line change
Expand Up @@ -268,7 +268,7 @@ inputOrgId })`. It resolves from: route param slug → request

**BYOK carve-out (intentional).** When a contributor sets a `BYOK_*`
key ([editor/lib/ai/models.ts](editor/lib/ai/models.ts) —
`BYOK_OPENROUTER_API_KEY`, `BYOK_AI_GATEWAY_API_KEY`), `grida`/`model`
`BYOK_OPENROUTER_API_KEY`, `BYOK_VERCEL_AI_GATEWAY_API_KEY`), `grida`/`model`
return a **bare** provider so the **AI-SDK text/chat path** bypasses
the billing seam: no gate, no Metronome ingest, **and** the
`MissingOrgIdError` runtime contract above does not fire (a bare
Expand All @@ -284,21 +284,35 @@ actions still read the real balance and cannot silently drain credit
while reporting `0`. **BYOK bypasses billing only — never auth.** `requireOrganizationId` and
route/action auth always run, so a logged-in user with no resolvable
org is still rejected. Gated solely by server-only, non-`NEXT_PUBLIC_`
env vars never set in the hosted product (same trust model as
`OPENAI_API_KEY` / `REPLICATE_API_TOKEN`). Fail-closed: `byok` is
env vars never set in the hosted product (same trust model as other
server-only credentials). Fail-closed: `byok` is
`null` unless a key env var is a non-empty string, so any ambiguity
falls back to the billed path. **Residual risk:** `byok` is resolved
once at module load with no per-request guard — an accidental `BYOK_*`
on a hosted/preview deploy would make every org bypass billing and the
org-id sanity gate (auth still holds). Acceptable only because it is a
contributor/self-host switch under the existing server-env trust model.

**Funded provider authority.** The shared Vercel AI Gateway provider reads only
`GG_VERCEL_AI_GATEWAY_API_KEY` for an explicit API key. An absent or blank value
retains the SDK's per-request platform OIDC resolution; an explicit empty SDK
setting prevents its ambient `AI_GATEWAY_API_KEY` lookup. Replicate predictions
require a nonblank `GG_REPLICATE_API_TOKEN` at execution, with no legacy token
fallback. Contributor overrides never supply these funded clients. Library query
embeddings still use the shared provider without user billing (or the contributor
override locally), and OpenAI model listing keeps its separate `OPENAI_API_KEY`.
Native CLI/Desktop credential names and stored provider IDs are independent of
this server configuration.

**Files bound by this id.** Run `grep -rn GRIDA-SEC-003 .` to enumerate.
Today:

- [editor/lib/auth/organization.ts](editor/lib/auth/organization.ts) — `requireOrganizationId`.
- [editor/lib/ai/server.ts](editor/lib/ai/server.ts) — single seam entry; unconditional runtime gate; BYOK layer switch.
- [editor/lib/ai/models.ts](editor/lib/ai/models.ts) — BYOK layer (bare provider, bypasses billing).
- [editor/lib/ai/models.ts](editor/lib/ai/models.ts) — contributor BYOK and explicit funded Vercel AI Gateway/OIDC authority.
- [Provider credential tests](editor/lib/ai/__tests__/models.test.ts) and
[Replicate credential tests](editor/lib/ai/__tests__/replicate-credentials.test.ts) —
funding-role separation, per-request OIDC, missing credentials and preserved billing.
- [GG Tripo execution](editor/lib/ai/gg-three-d.ts) and
[billing contract tests](editor/lib/ai/__tests__/gg-three-d.test.ts) — verified
org input, unconditional gate, infrastructure-key-only execution and actual receipts.
Expand Down Expand Up @@ -2507,7 +2521,7 @@ generation receipts without a fixed projection.
All selected file/environment/stdin keys use the shared AI provider admission
policy: bounded to 4 KiB, normalized, header-safe, free of known template values,
and checked against documented first-party formats without guessed suffix lengths.
Vercel legacy keys remain opaque where upstream specifies no retirement contract.
Vercel AI Gateway legacy keys remain opaque where upstream specifies no retirement contract.
Stdin has a cancellable 30-second bound. Only the
trusted SDK key reader receives secret strings. Status projects presence and
source and plaintext storage mode; disposal drops private references. Explicit
Expand Down Expand Up @@ -2566,7 +2580,7 @@ generation receipts without a fixed projection.
6. **Explicit credential checks before registration.** CLI `providers configure`
validates the entered key through the shared AI owner, then invokes that owner's
single authenticated GET before opening custody. The host permits only OpenRouter's
`/api/v1/key`, Vercel's `/v1/credits`, fal's `/v1/models/pricing` with exactly
`/api/v1/key`, Vercel AI Gateway's `/v1/credits`, fal's `/v1/models/pricing` with exactly
one fixed `endpoint_id=fal-ai/flux/dev`, and Tripo's `/v3/account/balance`.
This does not grant other platform APIs.
The shared owner rejects redirects, bounds the request/body lifecycle to ten seconds
Expand Down
2 changes: 1 addition & 1 deletion docs/cli/providers.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@ grida providers list
```

Enter your key at the hidden prompt. Every selected key passes static validation.
Configuration additionally checks OpenRouter, Vercel, and fal once before saving.
Configuration additionally checks OpenRouter, Vercel AI Gateway, and fal once before saving.
A rejected or unavailable check leaves existing credentials unchanged.
ElevenLabs has no suitable permission-neutral check and saves with verification
marked `not_supported`.
Expand Down
73 changes: 69 additions & 4 deletions docs/contributing/billing.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,7 @@
---
title: Contributing to Grida billing
description: Configure local billing sandboxes, contributor BYOK, and server provider credentials for Grida Gateway.
keywords: [grida, contributing, billing, credentials, byok, grida gateway]
format: md
---

Expand All @@ -21,18 +24,80 @@ If you are **not** working on the billing surface and only need the **AI chat /
# editor/.env.local (gitignored)
BYOK_OPENROUTER_API_KEY=sk-or-v1-... # https://openrouter.ai/keys
# …or, if you have one, a dedicated Vercel AI Gateway key:
# BYOK_AI_GATEWAY_API_KEY=...
# BYOK_VERCEL_AI_GATEWAY_API_KEY=...
```

- Bypasses **billing only — never auth.** Still sign in (`insider@grida.co` / `password`); a resolvable org is still required (an unauthenticated request still 401s).
- **Text/chat only** — BYOK swaps the AI-SDK provider, so only the text path is unbilled. Image/audio go through Replicate (`withTransaction`) and **still gate + bill even under BYOK** — those features need the full billing setup. (OpenRouter also exposes no image/audio models.) Catalog model IDs are unchanged; use IDs your provider accepts (edit `editor/lib/ai/models.ts` locally if one 404s).
- Precedence if both are set: OpenRouter, then Vercel. Fail-closed — an empty/unset (or whitespace-only) key falls back to the billed path.
- **Text/chat only** — the contributor override swaps the text provider. Hosted image, video, audio, and 3D operations **still gate + bill even under contributor BYOK** and need the full billing setup. Catalog model IDs are unchanged; use IDs your provider accepts.
- Precedence if both are set: OpenRouter, then Vercel AI Gateway. An empty/unset (or whitespace-only) contributor key does not select BYOK; without another contributor key, text uses the billed path.
- **Never set `BYOK_*` on a hosted or preview deploy.** It disables billing **and** the org-id sanity gate for every org. Contributor / self-host / local only. See [SECURITY.md](https://github.com/gridaco/grida/blob/main/SECURITY.md) (`GRIDA-SEC-003`, BYOK carve-out).

Working on billing itself? Ignore BYOK and continue with the full setup below.

---

## Server provider credentials

`GG_` identifies Grida-managed provider credentials used by the funded server
path. `BYOK_` identifies a contributor's server override. Neither prefix
identifies a deployment environment: scope secrets separately for Production,
Preview, and Development.

| Consumer | Configuration | Selection |
| ----------------------------------------------- | ---------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Funded Vercel AI Gateway text, image, and video | `GG_VERCEL_AI_GATEWAY_API_KEY` or platform `VERCEL_OIDC_TOKEN` | Use a nonblank GG key when configured; an unset or whitespace-only value selects platform OIDC. No fallback to `AI_GATEWAY_API_KEY` or contributor keys for funded authority. |
| Funded Replicate operations | `GG_REPLICATE_API_TOKEN` | Required and nonblank when a Replicate operation runs; no fallback to `REPLICATE_API_TOKEN` or BYOK. |
| Funded Tripo operations | `GG_TRIPO_API_KEY` | Unchanged; no generic or BYOK fallback. |
| Contributor text override | `BYOK_OPENROUTER_API_KEY`, then `BYOK_VERCEL_AI_GATEWAY_API_KEY` | Nonblank keys select the contributor's provider, bypassing billing only. |
| Library query embeddings | Shared contributor or Vercel AI Gateway provider | Remain unbilled internal operations with the existing contributor precedence. Sharing a funded provider credential does not add customer metering. |
| OpenAI model discovery | `OPENAI_API_KEY` | Unchanged; the model-list operation is nonbillable. |

Vercel supplies `VERCEL_OIDC_TOKEN` through its platform authentication lifecycle.
Keep that variable and its lifecycle intact; an OIDC deployment does not need a
new static Vercel AI Gateway key for the naming convention. Never copy an OIDC
token into a static secret. Missing provider configuration must fail at the
affected operation without preventing unrelated routes from loading.

These server names are separate from installed CLI/Desktop provider credentials.
The native provider ID remains `vercel`, and the CLI environment override remains
`AI_GATEWAY_API_KEY`. See [CLI provider keys](../cli/providers.md). The contributor
rename is a direct cutover: replace `BYOK_AI_GATEWAY_API_KEY` with
`BYOK_VERCEL_AI_GATEWAY_API_KEY` in local server configuration; the old name has
no compatibility alias.

### Deploying the credential rename

Code review and infrastructure changes are separate steps. The naming change
does not require rotating provider keys.

1. **A — code and draft PR.** Review the reader/consumer mapping, environment
examples, build environment forwarding, and synthetic credential-selection
tests. Record the infrastructure prerequisite in the draft PR; local work
does not authorize deployment or secret changes.
2. **B1 — before merge, after explicit operator GO.** Inventory environment scopes,
shared variables, and branch overrides. Add `GG_REPLICATE_API_TOKEN` in each
scope that runs funded Replicate operations, retaining `REPLICATE_API_TOKEN`
for the running deployment and rollback. If a deployment already uses a
Grida-managed `AI_GATEWAY_API_KEY`, add its value as
`GG_VERCEL_AI_GATEWAY_API_KEY` before deploying the new reader. OIDC deployments
need no Vercel AI Gateway API-key addition. Verify the candidate's affected provider
execution, billing, and shared Library embeddings before merge.
3. **B2 — after merge, after explicit operator GO.** Verify the intended deployed
revision and its effective configuration. After the agreed rollback window,
check for remaining consumers and remove obsolete server variables only from
migrated scopes. Retain any old key still needed by a rollback deployment or
another consumer. Removing an environment entry does not revoke the provider
key; revocation needs its own consumer check.

Treat secrets as values to transfer directly between approved stores, never as
review evidence. Record names and scopes without printing values. For write-only
secrets, an operator must enter the original value or issue a replacement while
retaining the old key through deployment verification. Validate the configuration
on a new deployment; editing project variables does not update an already
running deployment.

---

## What you need

- Local Supabase running (`supabase start`).
Expand Down Expand Up @@ -215,4 +280,4 @@ User-facing billing copy: [`docs/platform/billing.mdx`](../platform/billing.mdx)
| `WEBHOOK_TUNNEL_HOSTNAME` | `editor/.env.test.local` |
| `BILLING_E2E`, `BILLING_TEST_MODE`, `APP_URL` | `editor/.env.test` (committed) |

**Contributor BYOK (alternative — not required):** `BYOK_OPENROUTER_API_KEY` or `BYOK_AI_GATEWAY_API_KEY` in `editor/.env.local`. When set, the AI seam bypasses billing entirely and **none** of the Metronome rows above are needed. Auth is still required. See [Just need AI to work?](#just-need-ai-to-work-byok-instead-no-billing-setup).
**Contributor BYOK (alternative — not required):** `BYOK_OPENROUTER_API_KEY` or `BYOK_VERCEL_AI_GATEWAY_API_KEY` in `editor/.env.local`. When set, text/chat bypasses billing and **none** of the Metronome rows above are needed for that path. Auth is still required. See [Just need AI to work?](#just-need-ai-to-work-byok-instead-no-billing-setup).
2 changes: 1 addition & 1 deletion docs/editor/desktop/chatgpt-subscription.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ The model picker organizes other text models by provider:

- **Grida** is always shown. Its hosted models are metered against your
organization's prepaid Grida AI credit.
- **OpenRouter** and **Vercel** appear after their keys are configured.
- **OpenRouter** and **Vercel AI Gateway** appear after their keys are configured.
- **Ollama** appears after it is configured with at least one model.

Other provider groups appear when they are configured and available.
Expand Down
2 changes: 1 addition & 1 deletion docs/editor/desktop/local-models.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ machine, served by [Ollama](https://ollama.com). There is no account to
create and no API key to paste — your prompts, files, and the model's
responses never leave your computer.

You can use local models alongside provider keys (OpenRouter, Vercel), or
You can use local models alongside provider keys (OpenRouter, Vercel AI Gateway), or
as your only setup.

## Requirements
Expand Down
10 changes: 5 additions & 5 deletions docs/models/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -154,7 +154,7 @@ than applying GPT Image 2's per-image table. Its estimate excludes input
charges; use actual provider usage when comparing total costs. fal rounds
the total charge up to the nearest $0.0001.

[Vercel AI Gateway](https://vercel.com/ai-gateway/models/gpt-image-2.5-flare) publishes $5/M input tokens, $1.25/M cached input tokens, and $30/M output tokens. [OpenRouter](https://openrouter.ai/api/v1/images/models/openai/gpt-image-2.5-flare/endpoints) publishes $5/M text input, $8/M image input, and $30/M image output tokens. Grida keeps each provider's published meter separately. Hosted image billing uses Gateway's reported response cost when available. Without that receipt, it falls back to the catalog's coarse per-image estimate ($0.055 for these variants), which can differ from actual token cost across quality levels and dimensions.
[Vercel AI Gateway](https://vercel.com/ai-gateway/models/gpt-image-2.5-flare) publishes $5/M input tokens, $1.25/M cached input tokens, and $30/M output tokens. [OpenRouter](https://openrouter.ai/api/v1/images/models/openai/gpt-image-2.5-flare/endpoints) publishes $5/M text input, $8/M image input, and $30/M image output tokens. Grida keeps each provider's published meter separately. Hosted image billing uses Vercel AI Gateway's reported response cost when available. Without that receipt, it falls back to the catalog's coarse per-image estimate ($0.055 for these variants), which can differ from actual token cost across quality levels and dimensions.

**GPT Image 2** (`openai/gpt-image-2`) — _deprecated in Grida, superseded by GPT Image 2.5_

Expand Down Expand Up @@ -187,7 +187,7 @@ Provider support is checked separately from the model's native capability.
OpenAI added GPT Image 2 transparency in its [August 20 update](https://developers.openai.com/api/docs/changelog),
and fal exposes it. OpenRouter's [GPT Image 2 endpoint](https://openrouter.ai/api/v1/images/models/openai/gpt-image-2/endpoints)
and current GPT Image 2.5 endpoints still accept only `auto` or `opaque`.
Vercel transparency for GPT Image 2 remains unverified, although its
Vercel AI Gateway transparency for GPT Image 2 remains unverified, although its
[Flare](https://vercel.com/ai-gateway/models/gpt-image-2.5-flare) and
[Sunburst](https://vercel.com/ai-gateway/models/gpt-image-2.5-sunburst) pages
explicitly support it. Grida only offers transparency on verified routes.
Expand Down Expand Up @@ -305,8 +305,8 @@ which is cheaper on every provider; this is not an upstream Recraft retirement.
## Video Generation Models

Video models are billed per second of generated output, by resolution and
whether audio is generated. The rates below are the Grida-hosted (Vercel
gateway) rates; the hosted route always generates the model's default audio
whether audio is generated. The rates below are the Grida-hosted (Vercel AI
Gateway) rates; the hosted route always generates the model's default audio
mode, so the silent rates are informational until the request can carry an
audio mode.

Expand All @@ -323,7 +323,7 @@ provider meters one. Wan and Grok bundle audio into a single rate.

`Seedance 2.0` and `Seedance 2.5` (`bytedance/seedance-2.0`, `-2.5`) are
catalogued for bring-your-own-key use through fal, but are **not available on
the hosted route**: the gateway meters them per video token rather than per
the hosted route**: Vercel AI Gateway meters them per video token rather than per
second, and there is no honest per-second conversion, so Grida cannot
pre-price a hosted request. This will change when hosted video is metered
after generation. `Seedance 2.5` is the newer generation but not a cheaper
Expand Down
2 changes: 1 addition & 1 deletion docs/wg/ai/agent/chatgpt-subscription-provider.md
Original file line number Diff line number Diff line change
Expand Up @@ -479,7 +479,7 @@ the highest-capability compatible model (`GPT-5.6 Sol`). This default is derived
from live readiness and MUST NOT be persisted as a global preference. The
provider group order is ChatGPT Subscription when ready, Grida, configured text
BYOK providers, then configured compatible endpoints. The Grida, OpenRouter,
and Vercel groups may repeat catalog model ids because each tuple names a
and Vercel AI Gateway groups may repeat catalog model ids because each tuple names a
different cost and privacy boundary. OpenAI-shaped model ids outside the closed
ChatGPT set never gain ChatGPT eligibility from their name alone.

Expand Down
2 changes: 1 addition & 1 deletion docs/wg/cli/credential-custody.md
Original file line number Diff line number Diff line change
Expand Up @@ -166,7 +166,7 @@ claims about a provider's key length. The
[shared provider policy](https://github.com/gridaco/grida/blob/main/packages/grida-ai/README.md)
owns the exact rules, upstream references and authenticated check endpoints.

`providers configure` checks a newly entered OpenRouter, Vercel or fal key once
`providers configure` checks a newly entered OpenRouter, Vercel AI Gateway or fal key once
before saving, including when input comes from `--key-stdin`. A rejected,
permission-denied, timed-out or inconclusive check does not replace the old key.
ElevenLabs has no suitable permission-neutral check; its key is saved with static
Expand Down
2 changes: 1 addition & 1 deletion docs/wg/platform/billing/known-issues.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,7 +160,7 @@ costs. A single average invocation estimate cannot represent every request;
aggregate input/output token counts also cannot distinguish differently priced
text and image tokens.

**Current behavior.** When Gateway supplies a valid response cost, hosted image
**Current behavior.** When Vercel AI Gateway supplies a valid response cost, hosted image
usage is metered from that USD receipt, including a legitimate zero. Otherwise,
the existing catalog estimate is used. For GPT Image 2.5 that fallback is
$0.055 per requested image, regardless of quality and dimensions, and can
Expand Down
Loading
Loading