diff --git a/api-reference/admin-api/data-plane/logs/log-exports-beta/delete-a-log-export.mdx b/api-reference/admin-api/data-plane/logs/log-exports-beta/delete-a-log-export.mdx new file mode 100644 index 000000000..b5cd45435 --- /dev/null +++ b/api-reference/admin-api/data-plane/logs/log-exports-beta/delete-a-log-export.mdx @@ -0,0 +1,11 @@ +--- +title: Delete a Log Export +openapi: delete /logs/exports/{exportId} +--- + +Delete a completed, cancelled, or draft log export. In-progress exports must be cancelled before deletion. + + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; + + diff --git a/changelog/backend.mdx b/changelog/backend.mdx index f8ef9ce7b..b6ab11c0a 100644 --- a/changelog/backend.mdx +++ b/changelog/backend.mdx @@ -1,6 +1,6 @@ --- title: "Backend" -sidebarTitle: "Backend [1.22.0]" +sidebarTitle: "Backend [1.26.0]" rss: true --- @@ -8,6 +8,212 @@ rss: true Discuss how Portkey's AI Gateway can enhance your organization's AI infrastructure + + +## v1.26.0 +--- + +### New Providers +- **ElevenLabs**: Added to the Model Catalog (`elevenlabs`) for text-to-speech, speech-to-text, dubbing, and voice cloning + - [Documentation](/integrations/llms/elevenlabs) + +### JWT & Authentication +- **Multiple JWKS URL support**: JWT validation now accepts comma-separated JWKS URIs, allowing tokens from multiple identity providers to be verified in a single configuration. Keys from all endpoints are fetched in parallel and merged + - [Documentation](/product/mcp-gateway/authentication/jwt#multiple-jwks-uris) +- **Local JWT auth improvements**: Fixed JWT-authenticated requests from the gateway post-authentication, including proper workspace policy resolution and user details sync +- **JWKS cache optimization**: Improved key refetch logic with cooldown tracking to prevent redundant fetches, and fixed stale memory cache issues + +### Guardrails Updates +- **Webhook `executeOnProxy` parameter**: Webhook guardrails can now opt-in to run on proxy (`/v1/*`) pass-through requests via the `executeOnProxy` parameter, receiving method, path, and content-type metadata + - [Documentation](/integrations/guardrails/bring-your-own-guardrails#running-on-proxy-requests) + +### Log Exports +- **Delete log exports**: Log exports can now be deleted via `DELETE /v1/logs/exports/{id}`. In-progress exports must be cancelled before deletion. Deleted exports are archived and hidden from list/get APIs + - [Documentation](/product/observability/logs-export) + +### Performance & Infrastructure +- **AIRS timeout**: Default timeout field added for AIRS service requests +- **SCM config proxy**: Model fetching for SCM now routes through the config proxy +- **Organisation sync API**: Now returns usage policy IDs in organisation data responses + +### Security +- **CORS OPTIONS fix**: Fixed auth skip for OPTIONS preflight requests +- **Name sanitization**: Names and descriptions from the control plane are now sanitized before forwarding to the gateway, preventing encoding errors from special characters + +### Fixes and Improvements +- **Response handling**: Fixed unhandled error scenarios in response status handling +- **Large payload crash**: Fixed application crash when handling large response payloads +- **Workspace policies**: JWT auth response now correctly includes workspace-level usage policies + + + + + +## v1.25.0 +--- + +### MCP Gateway — Guardrails for MCP Servers +- MCP servers can now have **guardrails** attached via `capability_ids` mappings, with support for `run_on` values (`beforeToolCall`, `afterToolCall`). Guardrails are evaluated per tool call and can block or monitor tool execution +- New APIs for MCP server guardrail mappings with cache invalidation and lifecycle management +- MCP guardrail schemas added to the guardrail schema registry +- Guardrail list API no longer applies a default target filter, returning both LLM and MCP guardrails + +UI support for configuring MCP server guardrails is coming soon. + +### MCP Registry +- New MCP registry endpoint for listing available MCP servers with cursor-based pagination + +### New Guardrails +- **Singulr**: New partner guardrail integration (`singulr`) for AI security scanning +- **Alibaba Cloud AI Guardrails**: New partner guardrail integration for Alibaba Cloud's content moderation and security service + +### New Providers +- **Deepgram**: Added to the Model Catalog (`deepgram`) with TTS models and pricing +- **Lightning AI**: Added to the Model Catalog (`lightning-ai`) +- **Headroom**: Added to the Model Catalog (`headroom`) for enterprise AI integrations + +### Guardrails Updates +- **F5 Guardrails**: Added `forwardMetadata` parameter support +- **Config validation**: Guardrail slugs referenced in configs are now validated at save time — invalid or non-existent slugs in `input_guardrails` / `output_guardrails` are rejected +- Guardrail create API now returns `created_at` and `created_by` in the response + +### Integrations — Workload Identity Federation for Providers +- Integrations now support **workload identity federation** (`wif`) as a provider authentication type, enabling keyless authentication from cloud workloads (GKE, EKS, etc.) to AI providers + +### Deployment Enhancements +- Deployments support **tags** (key-value pairs) for filtering and organisation +- Deployments support **`source_type`** field and MCP Gateway base URL configuration +- List deployments API returns gateway base URLs + +### Observability +- Generation log responses now include **hook check data** (guardrail check results and errors) from the log store, surfacing detailed guardrail evaluation data alongside request/response logs +- Integration model auto-sync for automated model catalog refresh + +### Log Exports +- Log export deletion support +- Improved error messages when workspace is missing in export API requests + +### API Key Management +- User API key creation now rejects over-privileged scope requests — users cannot create keys with scopes exceeding their own role permissions +- `GET` workspace members API now returns user email + +### Self-Hosting +- Plugin schema now supports `service-role` for Bedrock plugins + +### Performance +- Read-through caching for organisation lookups reduces redundant DB queries +- Custom validator for `pricingConfig` validation on integration model updates + +### Security +- CORS headers excluded from prompt playground proxy forwarding +- Alibaba plugin credentials masked in API responses +- API key scope enforcement hardened for user role permissions +- Dependency updates and security fixes + +### Fixes and Improvements +- **Configs**: Fixed per-version author display in config version history (from v1.24.2 fix carry-over) +- **Prompts**: Fixed duplicate slug handling on prompt creation +- **Workspaces**: Fixed workspace slug hydrator validation error responses +- **Guardrails**: Fixed recursive guardrail slug extraction across nested config targets +- **Exports**: Fixed misleading workspace error message for exports API +- **Private Deployments**: Fixed virtual key handling scenarios in private deployments +- **SCIM**: Fixed group display name updates and SCIM group listing auto-map behavior +- **Cache**: Organisation cache improvements + + + + + +## v1.24.2 +--- + +### API Key Management — Default Scopes for Auto-Created User API Keys +- New **`user_api_key_auto_create_scopes`** organisation setting (with workspace-level override) controls which data-plane scopes are granted to auto-provisioned workspace-user API keys. When set, auto-created keys receive only the listed scopes instead of the full default scope set +- [Documentation](/product/administration/enforce-default-config#default-scopes-for-auto-created-user-api-keys) + +### JWT Authentication +- JWKS fetching now tries a direct fetch first and falls back to private link routing only when the direct path fails, improving resilience in hybrid deployments where the JWKS endpoint may or may not be reachable over the public internet + +### Guardrails +- `forwardHeaders` parameter removed from Portkey-managed PRO check schemas (`portkey.moderateContent`, `portkey.language`, `portkey.pii`, `portkey.gibberish`) — these checks are handled internally and do not support user-configured header forwarding + +### Security +- API key generation now uses strictly alphanumeric characters, eliminating modulo bias in the random generation + +### Fixes and Improvements +- **Configs**: Fixed per-version author and timestamp display in config version history + + + + + +## v1.24.1 +--- + +### Integrations — OAuth Client Credentials for Upstream Providers +- Integrations now support **`oauthClientCredentials`** as a `provider_auth_type`, enabling OAuth 2.0 Client Credentials grant for upstream provider authentication. Configure `oauth_token_endpoint`, `oauth_client_id`, `oauth_client_secret`, and optionally `oauth_scope` when creating or updating an integration +- Currently available for **OpenAI** integrations. The client secret is masked in API responses + +### Self-Hosting — Azure Workload Identity for Log Store +- Azure Blob Storage log store now supports **workload identity** authentication (`AZURE_AUTH_MODE=workload`), in addition to existing Entra ID and managed identity modes. Uses Kubernetes-projected federated tokens for keyless access to Azure Storage + +### Guardrails +- `forwardHeaders` parameter removed from 22 local-only guardrail check schemas (regex, model allowlist, word count, etc.) where it had no effect — only external checks that make outbound HTTP calls support header forwarding + + + + + +## v1.24.0 +--- + +### Rate Limit Policies — MCP Tool Call Target +- Rate limit policies now support a **`target`** field with values `llm` (default) and `mcp_tools`, allowing you to create rate limit policies specifically for MCP tool calls through the MCP Gateway +- The `target` field is returned in `GET` and `LIST` rate limit policy responses, and the `LIST` endpoint accepts a `target` query parameter for filtering +- MCP rate limits are included in the OAuth token introspection response for the MCP Gateway to enforce +- [Documentation](/product/enterprise-offering/budget-policies#target) + +Requires Gateway **2.18.0** or higher for MCP tool call rate limiting. + +### Guardrails — `forwardHeaders` Validation +- Guardrail `forwardHeaders` are now validated at create/update time. Restricted headers (credentials, cloud metadata, internal routing, hop-by-hop) and invalid header names are rejected +- The `forwardHeaders` parameter is now explicitly defined in all guardrail provider schemas for static schema export and gateway parity +- [Documentation](/product/guardrails/forwarding-headers) + +### Analytics — LLM, MCP, and A2A Breakdown +- Token, request, and error chart APIs now return separate counters for **LLM**, **MCP**, and **A2A** (agent-to-agent) traffic, enabling per-workload analytics breakdowns in the dashboard + +### Self-Hosting — Redis IAM & Workload Identity Authentication +- **AWS ElastiCache**: New IAM authentication mode (`AWS_REDIS_AUTH_MODE=iam`) for ElastiCache with automatic token refresh. Configure with `AWS_REDIS_CLUSTER_NAME`, `AWS_REDIS_REGION`, and optionally `AWS_REDIS_ASSUME_ROLE_ARN` / `AWS_REDIS_ROLE_EXTERNAL_ID` for cross-account access +- **Azure Cache for Redis**: Added workload identity authentication support via `AZURE_REDIS_WORKLOAD_CLIENT_ID` and `AZURE_REDIS_WORKLOAD_TENANT_ID`, and password authentication for Azure Redis +- **GCP Memorystore**: New `gcp-memory-store` cache store type with GCP workload identity authentication (`GCP_REDIS_AUTH_MODE`) +- [Components Documentation](/product/enterprise-offering/components) + +### Security +- Improved email masking in audit logs to preserve domain while distinguishing users + +### Fixes and Improvements +- **Prompts**: Fixed slug vs workspace ID resolution in prompt partials and prompts +- **Analytics**: Fixed cache status filter handling in cost and token chart APIs +- **Cache**: Fixed API key cache invalidation when organisation auth settings are updated +- Model configuration and pricing updates +- Security dependency updates + + + + + +## v1.23.0 +--- + +### Security +- Hardened SSRF checks for prompt playground and prompt completions proxy calls. Admin-supplied gateway URLs (from org enterprise settings or workspace deployments) are now routed through safe DNS lookup validation, blocking IMDS, metadata endpoints, and private DNS-rebind targets. Env-configured gateway hostnames are automatically added to the trusted host set + +### DB Migrations +- Added index on `api_keys.last_updated_at` for improved query performance + + + ## v1.22.0 diff --git a/changelog/enterprise.mdx b/changelog/enterprise.mdx index 807ff81f5..50350e451 100644 --- a/changelog/enterprise.mdx +++ b/changelog/enterprise.mdx @@ -1,6 +1,6 @@ --- title: "Enterprise Gateway" -sidebarTitle: "Enterprise Gateway [2.19.0]" +sidebarTitle: "Enterprise Gateway [2.21.0]" rss: true --- @@ -8,6 +8,116 @@ rss: true Discuss how Portkey's AI Gateway can enhance your organization's AI infrastructure + + +## v2.21.0 + +--- + +### ElevenLabs Provider + +ElevenLabs is now available as a provider, so text-to-speech and speech-to-text requests can route through the gateway with full observability and reliability features. Authentication uses ElevenLabs' `xi-api-key` header. Text-to-speech is billed on character count and speech-to-text on transcribed audio duration. + +[LLM Integrations](/integrations/llms) + +### Gateway-Local JWT Auth: Multiple JWKS URLs & User Attribution + +Gateway-local JWT authentication now accepts a comma-separated list of JWKS URLs, fetching and merging keys from all of them — useful when validating tokens issued by more than one identity provider. + +When a token's email resolves to an existing Portkey user in the target workspace, the request is now attributed to that user and inherits their workspace role, instead of being treated as a generic workspace service key. + +[JWT Authentication Documentation](/product/enterprise-offering/org-management/jwt) + +### Provider Updates + +- **SCX.ai**: Added as a new OpenAI-compatible provider for chat completions +- **Claude Platform on AWS**: Fixed the raw `anthropic-workspace-id` header (and `anthropic-beta`/`anthropic-version`) being dropped when requests use the structured `x-portkey-config` JSON shape instead of individual headers +- **Zhipu**: Usage is now passed through on streaming chat completions +- **Gemini**: Tool result messages are now mapped to the `user` role instead of `function`, matching Gemini's expected conversation format + +[Providers Documentation](/integrations/llms) + +### Fixes and Improvements + +- **JWT Auth**: Gateway-local JWT authentication now returns a clear 403 when a deployment restricts workspaces but none are configured, or when a resolved workspace can't be found, instead of silently falling back +- **Prisma AIRS Guardrails**: Scan requests now include `source_entity` and `source_entity_type` metadata for improved reporting, and fall back to the default timeout when an invalid (zero, negative, or non-numeric) custom timeout is configured +- **Security**: Updated dependencies to patch security vulnerabilities + + + + + +## v2.20.0 + +--- + +### MCP Gateway Guardrails + +MCP tool calls can now be checked against guardrails before and after execution, reusing the same guardrail checks available for LLM requests. Denied calls return a structured error with the guardrail and check ID for easy debugging. + +UI support is coming soon. + +[MCP Gateway Guardrails Documentation](/product/mcp-gateway/guardrails) + +### Workload Identity Federation for OpenAI & Anthropic + +OpenAI and Anthropic requests can now authenticate upstream using Workload Identity Federation, exchanging a workload identity token for a short-lived provider credential via OAuth client credentials instead of storing a static API key. + +[LLM Integrations](/integrations/llms) + +### xAI Image Generation + +The xAI provider now supports image generation, with OpenAI-compatible parameters plus xAI-specific aspect ratio and resolution controls. + +[xAI Documentation](/integrations/llms/x-ai) + +### Bedrock India Inference Profile & Context-Based Pricing + +Amazon Bedrock requests that use the India cross-region inference profile (model IDs prefixed with `in.`) now resolve to the correct underlying model. + +Bedrock models that charge different rates for short vs. long context are now billed at the tier matching the request's input token count, so long-context requests are no longer priced at the short-context rate. + +[AWS Bedrock Documentation](/integrations/llms/bedrock/aws-bedrock) + +### Unified Gateway + MCP Server Mode + +Set `SERVER_MODE=unified` to run the AI Gateway and MCP Gateway on a single port. In this mode, MCP endpoints are served under the `/m` path prefix (for example, `https://your-gateway.com/m/...`), while LLM traffic continues to use `/v1/*`. This removes the need to run two services or configure host-based routing to split AI Gateway and MCP Gateway traffic. + +[Self-Hosting Documentation](/self-hosting/hybrid-deployments/aws/eks) + +### Webhook Guardrails for Proxy Requests + +The webhook guardrail can now run on proxy (`/v1/*`) requests via an opt-in `executeOnProxy` parameter, letting you block passthrough traffic based on a custom webhook check. + +[Guardrails Documentation](/product/guardrails/list-of-guardrail-checks) + +### Provider Updates + +- **Azure AI Foundry**: Fixed cached input tokens being priced as regular input tokens for models called through `/v1/responses` +- **Azure OpenAI**: Codex models can now be called through the Anthropic-compatible `/v1/messages` endpoint +- **Azure OpenAI**: `entraFederated` and `workload` auth modes no longer send an empty `Authorization` header before token exchange completes +- **Amazon Bedrock**: Fixed a `temperature` validation error for Claude 4+ models called through the Messages API +- **Anthropic**: Long-running streams are no longer dropped during long model prefills — Anthropic's keepalive `ping` events are now passed through when streaming via the OpenAI-compatible `/v1/chat/completions` and `/v1/complete` endpoints +- **Anthropic**: Fixed prompt caching being silently skipped for messages with a single content block when calling Anthropic models through `/v1/responses` +- **fal.ai**: Fixed image-generation pricing falling back to the wrong model name +- **Together AI**: Added `stream_options` support +- **OpenAI**: Fixed a tool-call ID error when converting Anthropic tool calls to the Responses API + +[Providers Documentation](/integrations/llms) + +### Fixes and Improvements + +- **Security**: The JWT guardrail now fails closed by default on invalid signatures, `alg: none`, and expired tokens +- **Reliability**: Fixed `custom_host` leaking from one target to other targets in fallback/load-balance configs +- **JWT Validation**: Changes to your JWT auth configuration now take effect right away instead of waiting for previously cached validation results to expire +- **JWT Validation**: When a token introspection endpoint is configured with a custom `introspectContentType` (such as `application/json`), the request body is now encoded to match that content type instead of always being sent as form-urlencoded. Applies to both the JWT guardrail and MCP Gateway JWT authentication +- **MCP**: Removed a redundant per-user access check for MCP requests authenticated with local JWT auth. Local JWT mode does not carry user attribution, so there was no user to check access for +- **Usage Limits**: Fixed request-count-based usage limits and usage limits on JWT-authenticated keys not being applied — requests continued to be served after the configured limit was reached +- **Prisma AIRS Guardrails**: The scan endpoint is now configurable via the `AIRS_URL` environment variable, and scans now report the model, user, and provider from the live request instead of static config values +- **F5 Guardrails**: Added an opt-in `forwardMetadata` parameter to forward request metadata to the scan request + + + ## v2.19.0 @@ -42,7 +152,7 @@ The Qwen provider now supports the Responses and Messages endpoints, along with New Headroom plugin compresses request context before it reaches the model, reducing input token costs. Configured through the standard guardrails workflow. -[Guardrails Documentation](/product/guardrails) +[Headroom Plugin Documentation](https://portkey.ai/docs/integrations/guardrails/headroom) ### Provider-Reported Cost Billing diff --git a/changelog/frontend.mdx b/changelog/frontend.mdx index f8a46cabd..89f444efb 100644 --- a/changelog/frontend.mdx +++ b/changelog/frontend.mdx @@ -1,6 +1,6 @@ --- title: "Frontend" -sidebarTitle: "Frontend [1.9.4]" +sidebarTitle: "Frontend [1.9.6]" rss: true --- @@ -8,6 +8,76 @@ rss: true Discuss how Portkey's AI Gateway can enhance your organization's AI infrastructure + + +## v1.9.6 + +--- + +### Integrations + +- Added **Deepgram** as a new provider in the dashboard + +### Model Catalog + +- Added **OCR model type** support in the Model Catalog UI + +### Guardrails + +- Added **Singulr** guardrail partner logo and plugin UI +- Added **Headroom** plugin + +### Administration + +- Added **Owner role** support +- Added support for **configurable default user API key scopes** +- Added config to workspace update settings + +### Analytics + +- Added **tokens and request chart** token breakup API integration + +### Fixes and Improvements + +- **Configs**: Fixed `version_owner_id` to display the correct author per config version +- **Configs**: Fixed config create/edit to no longer block when guardrails API fails +- **Guardrails**: Fixed display of hook check data in Guardrails log view and resolved duplicate check ID rendering bug +- **Analytics**: Fixed chart time format issue +- **Target Policy**: Fixed target policy handling +- Audit fixes + + + + + +## v1.9.5 + +--- + +### Luna + +### Integrations + +- Added **Lightning AI** as a new provider in the dashboard + +### Model Catalog + +- Added support for **custom headers at model level** configuration + +### Bedrock + +- Updated **Bedrock API key** handling changes + +### Fixes and Improvements + +- **Analytics**: Fixed unique users chart to display higher-precision figures +- **MCP Gateway**: Fixed MCP redirect issue +- **OIDC**: Fixed OIDC redirection to preserve `organisation_id` +- **SAML**: Fixed SAML authentication flow +- Security vulnerability fixes + + + ## v1.9.4 diff --git a/docs.json b/docs.json index 3a40acf4b..b7f910c47 100644 --- a/docs.json +++ b/docs.json @@ -151,6 +151,7 @@ "pages": [ "product/mcp-gateway/authentication", "product/mcp-gateway/authentication/oauth", + "product/mcp-gateway/authentication/cas", "product/mcp-gateway/authentication/external-oauth", "product/mcp-gateway/authentication/identity-forwarding", "product/mcp-gateway/authentication/jwt", @@ -262,6 +263,12 @@ "product/enterprise-offering/org-management/scim/okta" ] }, + { + "group": "Directory Sync", + "pages": [ + "product/enterprise-offering/org-management/directory-sync/cie-directory-sync" + ] + }, "product/enterprise-offering/org-management/sso" ] }, @@ -462,6 +469,7 @@ "integrations/llms/deepbricks", "integrations/llms/deepgram", "integrations/llms/deepseek", + "integrations/llms/elevenlabs", "integrations/llms/github", "integrations/llms/groq", "integrations/llms/huggingface", @@ -490,6 +498,7 @@ "integrations/llms/reka-ai", "integrations/llms/recraft-ai", "integrations/llms/sambanova", + "integrations/llms/scalattice", "integrations/llms/segmind", "integrations/llms/snowflake-cortex", "integrations/llms/stability-ai", @@ -535,6 +544,7 @@ "integrations/guardrails/qualifire", "integrations/guardrails/javelin", "integrations/guardrails/f5-guardrails", + "integrations/guardrails/headroom", "integrations/guardrails/tavily", "integrations/guardrails/jwt", "integrations/guardrails/model-rules", @@ -605,6 +615,7 @@ "integrations/libraries/claude-desktop-developers" ] }, + "integrations/libraries/claude-for-microsoft-365", "integrations/libraries/claude-cowork", "integrations/libraries/opencode", "integrations/libraries/pi-agent", @@ -784,6 +795,7 @@ "api-reference/admin-api/data-plane/logs/log-exports-beta/create-a-log-export", "api-reference/admin-api/data-plane/logs/log-exports-beta/start-a-log-export", "api-reference/admin-api/data-plane/logs/log-exports-beta/cancel-a-log-export", + "api-reference/admin-api/data-plane/logs/log-exports-beta/delete-a-log-export", "api-reference/admin-api/data-plane/logs/log-exports-beta/download-a-log-export" ] } @@ -1939,6 +1951,7 @@ "api-reference/admin-api/data-plane/logs/log-exports-beta/create-a-log-export", "api-reference/admin-api/data-plane/logs/log-exports-beta/start-a-log-export", "api-reference/admin-api/data-plane/logs/log-exports-beta/cancel-a-log-export", + "api-reference/admin-api/data-plane/logs/log-exports-beta/delete-a-log-export", "api-reference/admin-api/data-plane/logs/log-exports-beta/download-a-log-export" ] } diff --git a/images/cas-consent.png b/images/cas-consent.png new file mode 100644 index 000000000..a298c2f3a Binary files /dev/null and b/images/cas-consent.png differ diff --git a/images/cas-login.png b/images/cas-login.png new file mode 100644 index 000000000..c70fb1bba Binary files /dev/null and b/images/cas-login.png differ diff --git a/images/directory-sync/add-group-mapping-dialog.png b/images/directory-sync/add-group-mapping-dialog.png new file mode 100644 index 000000000..9902ea317 Binary files /dev/null and b/images/directory-sync/add-group-mapping-dialog.png differ diff --git a/images/directory-sync/cie-directories-listing.png b/images/directory-sync/cie-directories-listing.png new file mode 100644 index 000000000..b4a7c5495 Binary files /dev/null and b/images/directory-sync/cie-directories-listing.png differ diff --git a/images/directory-sync/cie-set-up-directory.png b/images/directory-sync/cie-set-up-directory.png new file mode 100644 index 000000000..6067fdc73 Binary files /dev/null and b/images/directory-sync/cie-set-up-directory.png differ diff --git a/images/directory-sync/connected-directory-dropdown.png b/images/directory-sync/connected-directory-dropdown.png new file mode 100644 index 000000000..1fad2c9ad Binary files /dev/null and b/images/directory-sync/connected-directory-dropdown.png differ diff --git a/images/directory-sync/create-api-key-details.jpg b/images/directory-sync/create-api-key-details.jpg new file mode 100644 index 000000000..6b236df55 Binary files /dev/null and b/images/directory-sync/create-api-key-details.jpg differ diff --git a/images/directory-sync/create-api-key-filled.jpg b/images/directory-sync/create-api-key-filled.jpg new file mode 100644 index 000000000..e6bb6a8cb Binary files /dev/null and b/images/directory-sync/create-api-key-filled.jpg differ diff --git a/images/directory-sync/create-api-key-permissions.jpg b/images/directory-sync/create-api-key-permissions.jpg new file mode 100644 index 000000000..28897c01f Binary files /dev/null and b/images/directory-sync/create-api-key-permissions.jpg differ diff --git a/images/directory-sync/create-api-key-save.png b/images/directory-sync/create-api-key-save.png new file mode 100644 index 000000000..841562870 Binary files /dev/null and b/images/directory-sync/create-api-key-save.png differ diff --git a/images/directory-sync/create-api-key-select-user.png b/images/directory-sync/create-api-key-select-user.png new file mode 100644 index 000000000..107f602e8 Binary files /dev/null and b/images/directory-sync/create-api-key-select-user.png differ diff --git a/images/directory-sync/group-mappings.jpg b/images/directory-sync/group-mappings.jpg new file mode 100644 index 000000000..b8b2ca688 Binary files /dev/null and b/images/directory-sync/group-mappings.jpg differ diff --git a/images/directory-sync/portkey-directory-sync-overview.jpg b/images/directory-sync/portkey-directory-sync-overview.jpg new file mode 100644 index 000000000..0ad550426 Binary files /dev/null and b/images/directory-sync/portkey-directory-sync-overview.jpg differ diff --git a/images/directory-sync/security-keys-service-tab.jpg b/images/directory-sync/security-keys-service-tab.jpg new file mode 100644 index 000000000..01fd255e2 Binary files /dev/null and b/images/directory-sync/security-keys-service-tab.jpg differ diff --git a/images/directory-sync/security-keys-user-created.jpg b/images/directory-sync/security-keys-user-created.jpg new file mode 100644 index 000000000..9c9f0ea14 Binary files /dev/null and b/images/directory-sync/security-keys-user-created.jpg differ diff --git a/images/directory-sync/security-keys-user-tab.jpg b/images/directory-sync/security-keys-user-tab.jpg new file mode 100644 index 000000000..481a93ca5 Binary files /dev/null and b/images/directory-sync/security-keys-user-tab.jpg differ diff --git a/images/directory-sync/sync-state-success.jpg b/images/directory-sync/sync-state-success.jpg new file mode 100644 index 000000000..01609cca4 Binary files /dev/null and b/images/directory-sync/sync-state-success.jpg differ diff --git a/images/directory-sync/user-identity-attr-dropdown.png b/images/directory-sync/user-identity-attr-dropdown.png new file mode 100644 index 000000000..84bc6ebb9 Binary files /dev/null and b/images/directory-sync/user-identity-attr-dropdown.png differ diff --git a/images/directory-sync/workspace-control-list.jpg b/images/directory-sync/workspace-control-list.jpg new file mode 100644 index 000000000..3513cb1b1 Binary files /dev/null and b/images/directory-sync/workspace-control-list.jpg differ diff --git a/images/directory-sync/workspace-members.jpg b/images/directory-sync/workspace-members.jpg new file mode 100644 index 000000000..38b2f33bd Binary files /dev/null and b/images/directory-sync/workspace-members.jpg differ diff --git a/integrations/guardrails/bring-your-own-guardrails.mdx b/integrations/guardrails/bring-your-own-guardrails.mdx index f7ec3b3dc..f625a3803 100644 --- a/integrations/guardrails/bring-your-own-guardrails.mdx +++ b/integrations/guardrails/bring-your-own-guardrails.mdx @@ -32,10 +32,13 @@ In the Guardrail configuration UI, you'll need to provide: | **Webhook URL** | Your webhook's endpoint URL | `string` | | **Headers** | Headers to include with webhook requests | `JSON` | | **Timeout** | Maximum wait time for webhook response | `number` (ms) | +| **Fail on Error** | Treat a non-200 response or a timeout as a failed check | `boolean` | +| **Forward Headers** | Client headers to forward to your webhook | `string[]` | +| **Execute on Proxy** | Also run this webhook on proxy (`/v1/*`) requests | `boolean` | #### Webhook URL -This should be a publicly accessible URL where your webhook is hosted. +This should be a publicly accessible URL where your webhook is hosted. On self-hosted gateways, webhooks on private networks must be allowlisted in `TRUSTED_CUSTOM_HOSTS`. #### Headers @@ -56,6 +59,133 @@ The maximum time Portkey will wait for your webhook to respond before proceeding - Default: `3000ms` (3 seconds) - If your webhook processing is time-intensive, consider increasing this value +### Configure Your Webhook Inline in a Config + +Skip the portal entirely and pass the webhook check inline in the request config. Any hook that carries its own `checks` array runs as-is — no guardrail is created and no ID is looked up. + +```json Inline webhook guardrail +{ + "before_request_hooks": [{ + "id": "inline-input-webhook", + "type": "guardrail", + "deny": true, + "checks": [{ + "id": "default.webhook", + "parameters": { + "webhookURL": "https://guardrails.example.com/pre-check", + "headers": { "Authorization": "Bearer WEBHOOK_TOKEN" }, + "timeout": 5000, + "failOnError": true + } + }] + }], + "after_request_hooks": [{ + "id": "inline-output-webhook", + "type": "guardrail", + "deny": false, + "checks": [{ + "id": "default.webhook", + "parameters": { + "webhookURL": "https://guardrails.example.com/post-check", + "timeout": 3000 + } + }] + }] +} +``` + +Send the config with any inference request: + + +```sh cURL +curl https://api.portkey.ai/v1/chat/completions \ + -H "Content-Type: application/json" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -H 'x-portkey-config: {"before_request_hooks":[{"id":"inline-input-webhook","type":"guardrail","deny":true,"checks":[{"id":"default.webhook","parameters":{"webhookURL":"https://guardrails.example.com/pre-check","timeout":5000,"failOnError":true}}]}]}' \ + -d '{ + "model": "@my-openai/gpt-4o-mini", + "messages": [{"role": "user", "content": "Say Hi"}] + }' +``` + +```python Python +from portkey_ai import Portkey + +portkey = Portkey( + api_key="PORTKEY_API_KEY", + config={ + "before_request_hooks": [{ + "id": "inline-input-webhook", + "type": "guardrail", + "deny": True, + "checks": [{ + "id": "default.webhook", + "parameters": { + "webhookURL": "https://guardrails.example.com/pre-check", + "headers": {"Authorization": "Bearer WEBHOOK_TOKEN"}, + "timeout": 5000, + "failOnError": True + } + }] + }] + } +) + +response = portkey.chat.completions.create( + model="@my-openai/gpt-4o-mini", + messages=[{"role": "user", "content": "Say Hi"}] +) +``` + +```js NodeJS +import Portkey from 'portkey-ai' + +const portkey = new Portkey({ + apiKey: "PORTKEY_API_KEY", + config: { + before_request_hooks: [{ + id: "inline-input-webhook", + type: "guardrail", + deny: true, + checks: [{ + id: "default.webhook", + parameters: { + webhookURL: "https://guardrails.example.com/pre-check", + headers: { Authorization: "Bearer WEBHOOK_TOKEN" }, + timeout: 5000, + failOnError: true + } + }] + }] + } +}) + +const response = await portkey.chat.completions.create({ + model: "@my-openai/gpt-4o-mini", + messages: [{ role: "user", content: "Say Hi" }] +}) +``` + + +#### Hook Fields + +| Field | Description | Type | +| :-- | :-- | :-- | +| `id` | Label for the hook, shown in logs. Any string works — it is never looked up | `string` | +| `type` | Always `guardrail` for webhook checks | `string` | +| `checks` | The checks to run. One `default.webhook` entry per endpoint | `array` | +| `deny` | Block the request when the webhook returns `verdict: false` | `boolean` | +| `async` | Run the check without blocking the request | `boolean` | +| `sequential` | Run checks one after another instead of in parallel | `boolean` | + +To call several webhooks, add more entries to `checks`. They run in parallel unless `sequential` is `true`. + +Results come back under `hook_results.before_request_hooks` and `hook_results.after_request_hooks` in the response, exactly as they do for saved guardrails. + + +Inline checks still respect org-level guardrail settings — checks belonging to a disabled guardrail category are skipped. Configs travel in a request header, so very large inline configs can hit proxy header size limits. For big or shared guardrails, create a saved guardrail and reference it by ID instead. + + ### Webhook Request Structure Your webhook should accept `POST` requests with the following structure: @@ -416,6 +546,22 @@ Your webhook will receive these headers alongside `Content-Type` and any custom +## Running on Proxy Requests + +By default, webhook guardrails only run on managed `/v1/chat/completions`-style requests where Portkey parses the full request/response body. Enable `executeOnProxy` to also run your webhook on raw proxy (`/v1/*`) pass-through requests. + +```json +{ + "id": "default.webhook", + "parameters": { + "webhookURL": "https://my-guardrail.example.com/check", + "executeOnProxy": true + } +} +``` + +When triggered on a proxy request, the webhook payload includes `method`, `path`, and `content-type` metadata instead of a fully parsed request body. This is useful for blocking or auditing traffic that bypasses the standard chat completions route. + ## Important Implementation Notes 1. **Complete Transformations**: When using `transformedData`, include all fields in your transformed object, not just the changed portions. diff --git a/integrations/guardrails/headroom.mdx b/integrations/guardrails/headroom.mdx new file mode 100644 index 000000000..5fa3772e2 --- /dev/null +++ b/integrations/guardrails/headroom.mdx @@ -0,0 +1,155 @@ +--- +title: "Headroom" +description: "Headroom compresses LLM request context before inference to cut input token costs." +--- + +[Headroom](https://headroom-docs.vercel.app/) is a content-aware context compression engine. Portkey's Headroom guardrail sends the request's messages to a Headroom proxy, swaps in the compressed version, and forwards it to the model — so you pay for fewer input tokens without changing your application code. + + + + +Headroom is **bring-your-own-proxy**. Portkey does not host it. Deploy a Headroom proxy in your own infrastructure and point the integration at it — request payloads never leave your network. + + +## Deploy a Headroom Proxy + +```bash +pip install "headroom-ai[proxy]" + +HEADROOM_COMPRESS_ALLOW_REMOTE=1 \ +HEADROOM_PROXY_TOKEN="your-proxy-token" \ +headroom proxy --host 0.0.0.0 --port 8787 +``` + +| Variable | Purpose | +|----------|---------| +| `HEADROOM_COMPRESS_ALLOW_REMOTE=1` | Required when the gateway is not on the same host. Headroom's compression endpoint is loopback-only by default and returns `404` to every other caller. | +| `HEADROOM_PROXY_TOKEN` | Bearer token the proxy requires. Recommended once the proxy is reachable off-loopback. | + +Verify the deployment with `curl http://your-headroom-host:8787/health`. + +## Using Headroom with Portkey + +### 1. Add Headroom Credentials to Portkey + +* Click on the `Admin Settings` button on Sidebar +* Navigate to `Plugins` tab under Organisation Settings +* Click on the edit button for the **Headroom** integration +* Add your **Headroom Proxy URL** (for example `https://headroom.internal.example.com:8787`) +* Add your **Headroom Proxy Token** — leave blank if the proxy runs without `HEADROOM_PROXY_TOKEN` + +### 2. Add Headroom's Guardrail Check + +* Navigate to the `Guardrails` page and click the `Create` button +* Search for **"Headroom Compress Context"** and click `Add` +* Configure the compression behaviour (all parameters are optional): + +| Parameter | Type | Description | Default | +|-----------|------|-------------|---------| +| `mode` | enum | `optimize` sends the compressed payload to the provider. `audit` returns savings stats but forwards the original payload. | `optimize` | +| `compress_user_messages` | boolean | Compress user-role messages. Enable when user messages carry bulk data like logs or tool outputs. | `false` | +| `compress_system_messages` | boolean | Compress system-role messages. | `false` | +| `target_ratio` | number | Keep ratio for compression (`0.5` keeps 50%). Leave empty for Headroom defaults. | — | +| `protect_recent` | number | Number of recent messages left uncompressed. Set `0` to compress everything. | `4` | +| `protect_analysis_context` | boolean | Detect analyze/review intent and protect code from compression. | `true` | +| `token_budget` | number | Override the model's context limit. Headroom compresses to fit within this budget. | — | +| `timeout` | number | Maximum time to wait for the compression request (ms) | `10000` | + +* Set any `actions` you want on your check, and create the Guardrail! + + + Guardrail Actions allow you to orchestrate your guardrails logic. You can learn more about them [here](/product/guardrails#there-are-6-types-of-guardrail-actions) + + +| Check Name | Description | Supported Hooks | +|------------|-------------|-----------------| +| Headroom Compress Context | Compresses LLM request context using a Headroom proxy to reduce input token costs | `beforeRequestHook` | + +Start with `mode: audit` to measure savings on real traffic without changing what reaches the model, then switch to `optimize`. + +### 3. Add Guardrail ID to a Config and Make Your Request + +* When you save a Guardrail, you'll get an associated Guardrail ID — add this ID to the `input_guardrails` param in your Portkey Config +* Create these Configs in Portkey UI, save them, and get an associated Config ID to attach to your requests. [More here](/product/ai-gateway/configs). + +```json +{ + "input_guardrails": ["guardrails-id-xxx"] +} +``` + + + + +```js +const portkey = new Portkey({ + apiKey: "PORTKEY_API_KEY", + config: "pc-***" +}); +``` + + + +```py +portkey = Portkey( + api_key="PORTKEY_API_KEY", + config="pc-***" +) +``` + + + +```sh +curl https://api.portkey.ai/v1/chat/completions \ + -H "Content-Type: application/json" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -H "x-portkey-config: $CONFIG_ID" \ + -d '{ + "model": "gpt-4o-mini", + "messages": [{"role": "user", "content": "Analyse these logs: ..."}] + }' +``` + + + +For more, refer to the [Config documentation](/product/ai-gateway/configs). + +## Compression Results + +Savings appear in the `hook_results` block of the response and in your Portkey logs: + +```json +"data": { + "tokens_before": 12500, + "tokens_after": 4375, + "tokens_saved": 8125, + "compression_ratio": 0.35, + "transforms_applied": ["router:smart_crusher:0.35"], + "mode": "optimize" +} +``` + +`compression_ratio` is `tokens_after / tokens_before`, so lower is better — `0.35` means a 65% reduction. A ratio of `1.0` means nothing was compressed. + +## Compression Needs Large Inputs + +Headroom skips content where compression would cost more than it saves, so **small requests return zero savings by design**. Content below roughly 300 tokens, JSON arrays under 5 items, and tool outputs under 500 tokens pass through unchanged, as does source code and the last `protect_recent` messages. + +Headroom pays off on long agent sessions, multi-tool workflows, API and database responses, build output, and structured logs — typically 40–95% reduction. Short conversational turns see close to nothing. Test with a realistic payload, not a one-line prompt. + +## Supported Request Types + +Headroom compresses requests with a `messages` array: +- **Chat Completions** (`/v1/chat/completions`) +- **Messages** (Anthropic-style) + +Requests without a `messages` array pass through untouched. Compression also fails open — if the proxy is unreachable, slow, or errors, the original request is sent to the provider and the failure is recorded in `hook_results`. + +## Get Support + +If you face any issues with the Headroom integration, join the [Portkey community forum](https://discord.gg/portkey-llms-in-prod-1143393887742861333) for assistance. + + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; + + diff --git a/integrations/libraries/claude-for-microsoft-365.mdx b/integrations/libraries/claude-for-microsoft-365.mdx new file mode 100644 index 000000000..3054cc9a9 --- /dev/null +++ b/integrations/libraries/claude-for-microsoft-365.mdx @@ -0,0 +1,586 @@ +--- +title: "Claude for Microsoft 365" +description: "Deploy the Claude add-in for Excel, Word, PowerPoint, and Outlook through Portkey with governed model routing, scoped API keys, and full observability." +--- + +The Claude add-in for Microsoft 365 puts Claude in a task pane inside Excel, Word, PowerPoint, and Outlook. Point it at Portkey instead of Anthropic's API and every request from every user routes through one governed gateway. + +Configuration travels in the add-in's manifest XML, so nothing is left to individual users. Generate the manifest once, deploy it from M365 Admin Center, and the whole org is online. + +When you're done: + +- Every task-pane conversation routes through Portkey +- Model choice is pinned by config, not by the user +- Budget caps, rate limits, and guardrails enforce on every request +- Requests are attributed per user and team in Logs and Analytics + +```mermaid +flowchart LR + P["1.
Provider"] --> C["2.
Config"] + C --> K["3.
API Key"] + K --> V["4.
Verify"] + V --> M["5.
Manifest"] + M --> D["6.
Deploy"] + + style P fill:#ede9fe,stroke:#7c3aed,color:#1e1b4b + style C fill:#ddd6fe,stroke:#7c3aed,color:#1e1b4b + style K fill:#c4b5fd,stroke:#7c3aed,color:#1e1b4b + style V fill:#a78bfa,stroke:#7c3aed,color:#fff + style M fill:#8b5cf6,stroke:#7c3aed,color:#fff + style D fill:#7c3aed,stroke:#7c3aed,color:#fff +``` + +## How it works + +The add-in is hosted by Anthropic at `pivot.claude.ai`. You host nothing. The only artifact you produce is a manifest XML, and your gateway settings ride in it as URL query parameters on the task-pane URL. + +``` +Excel / Word / PowerPoint ──Authorization: Bearer──> Portkey ──> Anthropic / Bedrock / Vertex + │ + └── loads pivot.claude.ai/?gateway_url=…&gateway_token=…&gateway_api_format=anthropic + ▲ + manifest.xml +``` + +Model traffic goes from the user's Office WebView straight to Portkey. Allow `pivot.claude.ai` and your gateway host through the corporate firewall. + +## Prerequisites + +| Requirement | Notes | +|---|---| +| Portkey API key with a config attached | Steps 1–3 | +| Node.js 18+ | The manifest builder runs on Node | +| Claude Code | Hosts the manifest generator plugin | +| M365 Global Admin or Exchange Admin | Required to upload to Integrated apps | +| Excel/Word/PowerPoint 2019+ or Microsoft 365 | Windows, Mac, or web | + +## 1. Add a provider integration + +Connect the upstream model source Portkey calls on your behalf. + +Go to [Model Catalog](https://app.portkey.ai/model-catalog) → **Add Provider**. + + + + Direct API access + + + Cross-region inference + + + Google Cloud platform + + + +## 2. Create a config with `override_params` + +The add-in sends Anthropic-API model names from its built-in model picker — `claude-sonnet-4-5`, `claude-opus-4-5`, and so on. Those names won't resolve against your workspace's provider slugs, so pin the target model with `override_params`. + +Go to [Configs](https://app.portkey.ai/configs) → **Create Config**: + +```json +{ + "override_params": { + "model": "@anthropic-prod/claude-sonnet-4-20250514" + } +} +``` + +The value uses `@{provider-slug}/{model}` addressing, where the slug is the integration from step 1. + + + Naming convention: `m365-{team}-{env}`. Examples: `m365-finance-prod`, `m365-legal-prod`. + + +`override_params` **replaces** the matching field in the incoming request. Whatever model the add-in's picker sends is discarded, and every request is served by the model named here. + + + **The model picker becomes cosmetic.** Users switch models in the task pane and get identical behaviour. Collapse the picker to a single honest entry with [`available_models`](#optional-configuration), or route on the incoming model name instead of hard-pinning it. + + +Configs also carry reliability features. Open the relevant section if you need them: + + + + Route to a backup provider if the primary is unavailable: + + ```json + { + "strategy": { "mode": "fallback" }, + "targets": [ + { "provider": "@anthropic-prod" }, + { "provider": "@bedrock-prod" } + ] + } + ``` + + See [Fallbacks](/product/ai-gateway/fallbacks). + + + + Reduce cost and latency for repeated prompts: + + ```json + { + "provider": "@anthropic-prod", + "cache": { "mode": "simple" } + } + ``` + + See [Caching](/product/ai-gateway/cache-simple-and-semantic). + + + + Send different teams to different models using metadata from step 4: + + ```json + { + "strategy": { + "mode": "conditional", + "conditions": [ + { "query": { "metadata.team": { "$eq": "research" } }, "then": "opus-target" } + ], + "default": "sonnet-target" + }, + "targets": [ + { + "name": "opus-target", + "provider": "@anthropic-prod", + "override_params": { "model": "claude-opus-4-5" } + }, + { + "name": "sonnet-target", + "provider": "@anthropic-prod", + "override_params": { "model": "claude-sonnet-4-5" } + } + ] + } + ``` + + See [Conditional Routing](/product/ai-gateway/conditional-routing). + + + +## 3. Create the API key + +Go to [API Keys](https://app.portkey.ai/api-keys) → create one key per team or environment. + + + + Every request authenticated with this key inherits the config. + + + See [Budget Limits](/product/model-catalog/budget-limits) and [Rate Limits](/product/model-catalog/rate-limits). + + + Required. Without it the attached config does not apply to the request. + + + + + **Allow config override is not optional.** Left off, Portkey keeps the client's own parameters and passes the add-in's `claude-sonnet-4-5` straight through to the provider instead of replacing it with your `override_params.model`. The failure is quiet — a 200 from an unintended model, or a provider error about an unknown model, with nothing naming the config as the cause. + + +## 4. Verify the gateway before building anything + +A manifest pointing at an unreachable model deploys perfectly and then fails at the user's first message. Probe first. + +```sh +curl "https://api.portkey.ai/v1/messages" \ + -H 'content-type: application/json' \ + -H 'authorization: Bearer $PORTKEY_API_KEY' \ + -d '{"model": "claude-sonnet-4-5", "max_tokens": 1, "messages": [{"role": "user", "content": "hi"}]}' +``` + +Check two things in the response: + +| Check | Expected | +|---|---| +| HTTP status | `200` | +| `model` field | The model pinned in step 2 — **not** the model sent in the request | + +If the response echoes `claude-sonnet-4-5`, the override isn't applying. Re-check *Allow config override* on the API key. + + + **Use `authorization`, not `x-api-key`.** The add-in supports only those two header schemes, and Portkey rejects its API key on `x-api-key` with `Invalid API Key. Error Code: 03`. Portkey's native `x-portkey-api-key` works with curl but the add-in cannot send it. + + +## 5. Install the manifest generator + +The manifest is built by `claude-for-msft-365-install`, a Claude Code plugin published by Anthropic. It fetches the canonical manifest template and appends your gateway settings as URL query parameters. + + + + The build and validation scripts run on Node 18+. + + ```sh + node --version + ``` + + + ```sh + claude plugin marketplace add anthropics/financial-services + claude plugin install claude-for-msft-365-install@claude-for-financial-services + ``` + + + Slash commands register at startup. Without a restart they won't appear. + + + The plugin installs to a version-pinned directory: + + ```sh + PLUGIN=~/.claude/plugins/cache/claude-for-financial-services/claude-for-msft-365-install/0.1.11 + ``` + + Check your version with `ls ~/.claude/plugins/cache/claude-for-financial-services/claude-for-msft-365-install/`. Inside a session, `${CLAUDE_PLUGIN_ROOT}` resolves automatically. + + + +The plugin adds these commands: + +| Command | Purpose | +|---|---| +| `/claude-for-msft-365-install:setup` | Guided wizard — walks the full deployment end to end | +| `/claude-for-msft-365-install:manifest` | Generate the manifest XML only | +| `/claude-for-msft-365-install:consent` | Azure admin consent URLs for Entra SSO and Outlook Graph | +| `/claude-for-msft-365-install:bootstrap` | Build a per-user config endpoint | +| `/claude-for-msft-365-install:update-user-attrs` | Per-user config via Entra extension attributes | +| `/claude-for-msft-365-install:access-policies` | Scope features by Purview sensitivity label | +| `/claude-for-msft-365-install:debug` | Diagnose a broken deployment | +| `/claude-for-msft-365-install:export-data` | Export a user's chat history, skills, and MCP registrations | + +Update later with `claude plugin update claude-for-msft-365-install@claude-for-financial-services`, then restart. + +### Option A — guided wizard + +Run the wizard and answer its prompts. It handles everything from provider choice through Admin Center upload: + +```sh +/claude-for-msft-365-install:setup +``` + +| Wizard step | What it asks | +|---|---| +| 1. Connection path | Gateway vs. Vertex/Bedrock/Foundry. **Answer gateway** — Portkey is the gateway, even though it routes to Bedrock or Vertex underneath | +| 1b. Office apps | Excel/Word/PowerPoint, Outlook, or both | +| 2. Admin consent | Skipped unless `entra_sso=1` is needed | +| 3. Org-wide vs per-user | Whether the gateway token is the same for everyone | +| 4. Generate manifest | Runs the build script with your captured keys | +| 5. Per-user config | Extension attrs or a bootstrap endpoint, if step 3 needs them | +| 6. Verify a model | The probe from step 4 above | +| 7. Deploy | Walks the Admin Center upload | + + + The wizard appends every command and captured value to a setup log at `~/Desktop/claude-for-msft-365-install-setup.md`. Re-running it resumes from that log instead of starting over. + + +### Option B — generate directly + +To skip the wizard, call the build script yourself. Excel, Word, and PowerPoint share one manifest. Outlook uses a different Microsoft schema and needs its own: + + +```sh Excel / Word / PowerPoint +node "$PLUGIN/scripts/build-manifest.mjs" office manifest.xml \ + gateway_url=https://api.portkey.ai \ + gateway_token='$PORTKEY_API_KEY' \ + gateway_auth_header=authorization \ + gateway_api_format=anthropic +``` + +```sh Outlook +node "$PLUGIN/scripts/build-manifest.mjs" outlook manifest-outlook.xml \ + gateway_url=https://api.portkey.ai \ + gateway_token='$PORTKEY_API_KEY' \ + gateway_auth_header=authorization \ + gateway_api_format=anthropic \ + graph_client_id= +``` + + +### The four gateway keys + +| Key | Value for Portkey | +|---|---| +| `gateway_url` | `https://api.portkey.ai`, or your self-hosted gateway URL | +| `gateway_token` | The API key from step 3 | +| `gateway_auth_header` | `authorization` — the `x-api-key` default does not work | +| `gateway_api_format` | `anthropic` — Portkey exposes the `/v1/messages` route | + + + Single-quote the token in shell. Portkey keys contain `/` and `+`; the builder percent-encodes them correctly, but an unquoted shell mangles them first. + + +Validate before uploading: + +```sh +npx --yes office-addin-manifest validate manifest.xml +``` + + + **Outlook needs Microsoft Graph admin consent** before deployment, or every user hits "Need admin approval" on first open. Run `/claude-for-msft-365-install:consent` first. Note that Outlook does not support the Bedrock direct path. + + +### Version and Id + +```xml +4e2fd29d-dad0-40b3-af3c-a16e347a5ddc +1.0.0.12 +``` + +First deploy: leave both alone. Updating an existing deployment: bump the fourth segment of ``. M365 Admin Center caches by `` + `` and **silently ignores** a re-upload at the same version — the most common cause of "I updated it but nothing changed". + +## 6. Test locally, then deploy + +Sideload on one machine first. This bypasses the 24–72 hour Admin Center cache entirely, and a sideloaded manifest wins over a centrally deployed one with the same ``. + + +```sh macOS +bash "$PLUGIN/scripts/sideload-addin.sh" manifest.xml +``` + +```powershell Windows +& "$PLUGIN\scripts\sideload-addin.ps1" C:\path\to\manifest.xml +``` + + +Fully quit and reopen Excel — a backgrounded app won't rescan. The add-in appears under **Insert → My Add-ins**. Send a message and confirm the request lands in [Logs](https://app.portkey.ai/logs). + +Remove a sideload with `clear-addin-cache.sh --id --apply`. It's dry-run by default and only ever touches that one ID. + +### Deploy the add-in for your organization + + + + Go to [M365 Admin Center → Settings → Integrated apps](https://admin.cloud.microsoft/?#/Settings/IntegratedApps) and choose **Upload custom apps**. + + + App type: **Office Add-in**. Choose how to upload: **Upload manifest file (.xml) from device**, then select `manifest.xml`. + + Admin Center validates on upload — the same check as `office-addin-manifest validate`. + + + Start with **Just me** or a pilot group. Widen to **Entire organization** once verified — assignment changes don't require redeployment. + + Assign to **Specific users/groups** if per-user config was issued, matching exactly who was provisioned. Nested groups aren't supported. + + + Review the requested permissions, then **Finish deployment**. + + + Upload `manifest-outlook.xml` as a separate app if Outlook was generated. + + + +The add-in appears under **Home → Add-ins** in Excel, Word, and PowerPoint once it lands. + + + Propagation takes up to 24 hours for a fresh deploy and up to 72 hours for an update. To skip the update wait, redeploy with a fresh `` UUID — every client then treats it as a brand-new add-in. + + +## Update the manifest + +Rotating the Portkey key, changing the model list, or adding any config key means regenerating and re-uploading. Run the four commands in this order: + + +### 1. Regenerate with the new values + +```sh +node "$PLUGIN/scripts/build-manifest.mjs" office manifest.xml \ + gateway_url=https://api.portkey.ai \ + gateway_token='$NEW_PORTKEY_API_KEY' \ + gateway_auth_header=authorization \ + gateway_api_format=anthropic \ + available_models='claude-sonnet-4-5' +``` + +### 2. Bump the fourth version segment +```sh +perl -i -pe 's{(\d+\.\d+\.\d+\.)(\d+)()}{$1.($2+1).$3}e' manifest.xml +``` + +### 3. Confirm the bump and validate +```sh +grep -o "[^<]*" manifest.xml +npx --yes office-addin-manifest validate manifest.xml +``` +### 4. Re-upload in Admin Center → Integrated apps → your add-in → Update + + + **The build script overwrites the file.** It fetches the canonical template and writes it out fresh — it never reads your existing `manifest.xml`. Any hand-edit is discarded, including a `` bump, so always bump *after* regenerating, never before. + + +Hand-editing the manifest is otherwise fine, and those edits do reach users — trimming `` entries to drop PowerPoint, for example. Re-apply them after each regeneration, and bump `` regardless of how the change was made, or Admin Center serves the cached copy. + + + Keep `` unchanged so Admin Center treats this as an update to the existing add-in rather than a second, parallel installation. + + +To verify a change immediately instead of waiting out the cache, clear and re-sideload on one machine: + +```sh +bash "$PLUGIN/scripts/clear-addin-cache.sh" --id --apply +bash "$PLUGIN/scripts/sideload-addin.sh" manifest.xml +``` + +Fully quit and reopen the Office app — clearing does nothing until the app re-reads on launch, and a backgrounded app counts as still running. + +## Attribute requests per user and team + +Add `inference_headers` to the manifest to tag every request with metadata Portkey uses for filtering, cost attribution, and conditional routing: + +```sh +inference_headers='{"x-portkey-metadata": "{\"team\":\"finance\",\"env\":\"prod\",\"app\":\"m365-addin\"}"}' +``` + +| Use case | How it works | +|---|---| +| **Filter Logs** | Search by any field combination, e.g. `team=finance` and `env=prod` | +| **Attribute spend** | Group cost, tokens, and latency by team or environment in Analytics | +| **Conditional routing** | Route on metadata values in the step 2 config | + + + `Authorization`, `x-api-key`, `Content-Type`, `Host`, `Content-Length`, `User-Agent`, `Cookie`, and any `anthropic-*` / `x-amz-*` / `x-goog-*` header are reserved and silently dropped — they carry the add-in's own auth and protocol negotiation. + + +For per-user metadata rather than one org-wide value, issue narrower keys per team, or serve per-user config from a bootstrap endpoint (`/claude-for-msft-365-install:bootstrap`). + +## Optional configuration + +Each of these is another `key=value` argument to the build command. + + + + **Overrides** the picker. Users see exactly what's listed, in order, and nothing else. List every model that should remain, not just additions. + + ```sh + available_models='claude-sonnet-4-5' + available_models='[{"id": "claude-opus-4-5", "label": "Opus"}, "claude-sonnet-4-5"]' + ``` + + Set this to a single entry when the step 2 config pins one model, so the picker tells the truth. + + + + JSON array of MCP servers the add-in connects to directly. `headers` present means static auth; absent triggers OAuth discovery. Values interpolate other config keys. + + ```sh + mcp_servers='[{"url": "{{gateway_url}}/deepwiki/mcp", "label": "DeepWiki", "headers": {"Authorization": "Bearer {{gateway_token}}"}}]' + ``` + + See [MCP Gateway](/product/mcp-gateway). + + + + Comma-separated slugs in `{domain}.{action}` form: + + | Slug | Effect | + |---|---| + | `skills.authoring` | Blocks creating, editing, and uploading skills | + | `file.upload` | Blocks attaching files to the conversation | + | `web_search` | Removes native web search and fetch, which otherwise egress to a third-party provider | + | `thumbs` | Blocks response feedback | + | `addin.access` | Kill switch — the add-in refuses to run | + + Pair `web_search` with `mcp_servers` to substitute an in-network search tool. Unknown slugs are ignored. + + + + By default users skip the connection form and land straight in chat. Set `auto_connect=0` to show the form prefilled instead. + + The **Back** button to Claude.ai sign-in is hidden whenever enterprise config is present. Set `allow_1p=1` to keep it. + + + + Routes the add-in's OpenTelemetry traces to a collector you operate. Set the base HTTPS URL; the add-in appends `/v1/traces` and posts OTLP/HTTP. gRPC isn't supported — the add-in runs in a browser WebView. + + Portkey also exports traces natively. See [OpenTelemetry](/product/observability/opentelemetry). + + + +## Security + + + **The `gateway_token` in a manifest is a shared secret distributed to every user.** It sits in plaintext in the task-pane URL, readable from any installed machine's add-in cache. Don't commit a manifest containing a live key to git. + + +Mitigate at the Portkey layer rather than trying to hide the key: + +- Scope each key narrowly with its own config, budget cap, and rate limit +- Issue one key per team so revocation and rotation are surgical +- Rotate on a schedule — regenerating the manifest is one command + +For per-user tokens instead of one shared key, serve them from a bootstrap endpoint (`/claude-for-msft-365-install:bootstrap`), which returns per-user JSON config at startup and overrides manifest values. + +## Pre-launch checklist + + + + Send a test message from the task pane. Confirm it shows in [Logs](https://app.portkey.ai/logs) with the expected metadata. + + + Check the `model` field in the log entry. It should match the step 2 config, not the add-in's picker. + + + Test each budget cap, rate limit, and guardrail you configured. + + + Verify spend rolls up under the right team in Analytics. + + + `pivot.claude.ai` for the add-in payload, and your gateway host for model traffic. + + + +## Troubleshooting + + + + The config isn't applying. Confirm *Allow config override* is on for the API key (step 3) and that a config is attached to it. + + + + Wrong header scheme. The add-in must use `gateway_auth_header=authorization`; Portkey rejects its key on `x-api-key`. Re-run the step 4 probe. + + + + Two caches. Admin Center ignores re-uploads at the same `` — bump the fourth segment. Then the client Wef cache holds until the app restarts; clear it with `clear-addin-cache.sh --id --apply` and fully quit Office. Service-side propagation takes up to 72 hours for updates. + + + + Check Admin Center → Integrated apps → Users tab. Nested groups aren't supported. If it shows under My Add-ins but has no ribbon button, the manifest's `` is missing that app — check both the top-level `` list and the one under ``. + + + + The error screen has a **Copy error details** button. The paste shows a `Request:` block and a `Manifest params:` block with identical key names — diff them. Matching values mean the manifest went through unchanged and the problem is upstream. `Raw error:` is ground truth. + + + + **macOS:** quit the app, run `defaults write com.microsoft.Excel OfficeWebAddinDeveloperExtras -bool true`, enable Safari's developer features, then enable your terminal under System Settings → Privacy & Security → Developer Tools. That third gate is the one everyone misses. Right-click in the task pane → **Inspect Element**. + + **Windows:** right-click in the task pane → **Inspect**. No setup needed with WebView2. + + + + + Chat history, skills, and MCP registrations live in browser storage on the user's own machine — there is no server-side copy. On Windows the widely circulated "delete the `Wef` folder" fix destroys them along with the manifest cache. Export first with `/claude-for-msft-365-install:export-data`. + + +## Related + + + + Roll out Claude Desktop across your org + + + Route Claude Code through Portkey + + + Routing, fallbacks, and caching + + + + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; + + diff --git a/integrations/llms/elevenlabs.mdx b/integrations/llms/elevenlabs.mdx new file mode 100644 index 000000000..ed3295368 --- /dev/null +++ b/integrations/llms/elevenlabs.mdx @@ -0,0 +1,126 @@ +--- +title: "ElevenLabs" +description: Use ElevenLabs' text-to-speech, speech-to-text, and voice cloning APIs through Portkey. +--- + +ElevenLabs provides industry-leading text-to-speech, speech-to-text, dubbing, and voice cloning capabilities. Portkey allows you to use ElevenLabs' APIs with full observability and reliability features. + + +ElevenLabs currently uses a custom host setup. SDK support is available through the custom host pattern. + + +## Quick Start + +### Text to Speech + +Generate speech from text using the ElevenLabs TTS API: + +```bash cURL +curl -X POST https://api.portkey.ai/v1/audio/speech \ + -H "xi-api-key: $ELEVENLABS_API_KEY" \ + -H "Content-Type: application/json" \ + -H "x-portkey-custom-host: https://api.elevenlabs.io" \ + -H "x-portkey-provider: openai" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -d '{ + "model_id": "eleven_multilingual_v2", + "text": "Hello! Welcome to Portkey.", + "voice_settings": { + "stability": 0.5, + "similarity_boost": 0.75 + } + }' \ + --output speech.mp3 +``` + +## Add Provider in Model Catalog + +Before making requests, add ElevenLabs to your Model Catalog: + +1. Go to [**Model Catalog → Add Provider**](https://app.portkey.ai/model-catalog/providers) +2. Select **ElevenLabs** +3. Enter your [ElevenLabs API key](https://elevenlabs.io/app/settings/api-keys) +4. Set the custom host to `https://api.elevenlabs.io` +5. Name your provider (e.g., `elevenlabs`) + + + See all setup options and detailed configuration instructions + + +--- + +## Using with Portkey SDK + +You can also use ElevenLabs with the Portkey SDK using custom host configuration: + + + +```python Python +from portkey_ai import Portkey + +portkey = Portkey( + api_key="PORTKEY_API_KEY", + provider="openai", + custom_host="https://api.elevenlabs.io", + Authorization="Bearer ELEVENLABS_API_KEY" +) +``` + +```javascript Node.js +import Portkey from 'portkey-ai'; + +const portkey = new Portkey({ + apiKey: 'PORTKEY_API_KEY', + provider: 'openai', + customHost: 'https://api.elevenlabs.io', + Authorization: 'Bearer ELEVENLABS_API_KEY' +}); +``` + + + +--- + +## ElevenLabs Features + +ElevenLabs offers advanced audio AI capabilities: + +- **Text to Speech**: Convert text to natural-sounding speech in 29+ languages +- **Speech to Text**: Transcribe audio with high accuracy +- **Voice Cloning**: Clone voices from audio samples +- **Dubbing**: Automatically dub audio/video content across languages +- **Voice Design**: Generate unique AI voices + + + Explore the complete ElevenLabs API documentation + + +--- + +## Next Steps + + + + Add retries and timeouts for audio processing + + + Monitor and trace your ElevenLabs requests + + + Learn more about custom host configuration + + + Add custom metadata to track audio sources + + + +For complete SDK documentation: + + + Complete Portkey SDK documentation + + + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; + + diff --git a/integrations/llms/scalattice.mdx b/integrations/llms/scalattice.mdx new file mode 100644 index 000000000..98d8381c6 --- /dev/null +++ b/integrations/llms/scalattice.mdx @@ -0,0 +1,164 @@ +--- +title: "Scalattice" +description: Use Scalattice catalog models through Portkey's AI Gateway. +--- + +Portkey provides a robust and secure gateway to integrate Large Language Models (LLMs) into your applications, including [Scalattice](https://scalattice.cloud/docs/developers). + +With Portkey, you get fast AI gateway access, observability, and prompt management, while API keys stay in [Model Catalog](/product/model-catalog). + +Provider Slug. `scalattice` + +## Quick Start + +Get started with Scalattice in a couple of minutes: + + + +```python Python icon="python" +from portkey_ai import Portkey + +# 1. Install: pip install portkey-ai +# 2. Add @scalattice provider in model catalog +# 3. Use it: + +portkey = Portkey(api_key="PORTKEY_API_KEY") + +response = portkey.chat.completions.create( + model="@scalattice/qwen-3-14b", + messages=[{"role": "user", "content": "Hello!"}] +) + +print(response.choices[0].message.content) +``` + +```js Javascript icon="square-js" +import Portkey from 'portkey-ai' + +// 1. Install: npm install portkey-ai +// 2. Add @scalattice provider in model catalog +// 3. Use it: + +const portkey = new Portkey({ + apiKey: "PORTKEY_API_KEY" +}) + +const response = await portkey.chat.completions.create({ + model: "@scalattice/qwen-3-14b", + messages: [{ role: "user", content: "Hello!" }] +}) + +console.log(response.choices[0].message.content) +``` + +```python OpenAI Py icon="python" +from openai import OpenAI +from portkey_ai import PORTKEY_GATEWAY_URL + +# 1. Install: pip install openai portkey-ai +# 2. Add @scalattice provider in model catalog +# 3. Use it: + +client = OpenAI( + api_key="PORTKEY_API_KEY", + base_url=PORTKEY_GATEWAY_URL +) + +response = client.chat.completions.create( + model="@scalattice/qwen-3-14b", + messages=[{"role": "user", "content": "Hello!"}] +) + +print(response.choices[0].message.content) +``` + +```js OpenAI JS icon="square-js" +import OpenAI from "openai" +import { PORTKEY_GATEWAY_URL } from "portkey-ai" + +// 1. Install: npm install openai portkey-ai +// 2. Add @scalattice provider in model catalog +// 3. Use it: + +const client = new OpenAI({ + apiKey: "PORTKEY_API_KEY", + baseURL: PORTKEY_GATEWAY_URL +}) + +const response = await client.chat.completions.create({ + model: "@scalattice/qwen-3-14b", + messages: [{ role: "user", content: "Hello!" }] +}) + +console.log(response.choices[0].message.content) +``` + +```sh cURL icon="square-terminal" +# 1. Add @scalattice provider in model catalog +# 2. Use it: + +curl https://api.portkey.ai/v1/chat/completions \ + -H "Content-Type: application/json" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -d '{ + "model": "@scalattice/qwen-3-14b", + "messages": [{"role": "user", "content": "Hello!"}] + }' +``` + + + +## Add Provider in Model Catalog + +Before making requests, add Scalattice to your Model Catalog: + +1. Go to [**Model Catalog → Add Provider**](https://app.portkey.ai/model-catalog/providers) +2. Select **Scalattice** (or an OpenAI-compatible custom host if the named entry is not in the dropdown yet) +3. Enter your [Scalattice API key](https://scalattice.cloud/developers) (`slt_...`) +4. Name your provider (for example, `scalattice`) + + + See all setup options and detailed configuration instructions + + +Base URL: `https://api.scalattice.cloud/v1` + +## Chat Completions + +Scalattice is OpenAI-compatible for chat completions, including streaming and tools. Catalog IDs match `GET https://api.scalattice.cloud/v1/models`. + +Popular models: + +| Model | Notes | +|-------|-------| +| `qwen-3-14b` | Default mid-size general model | +| `qwen-3-coder-30b-a3b` | Long-context coding | +| `llama-3.3-70b` | High quality open chat | +| `qwen-3-vl-8b` | Vision + text | + +## Direct provider header + +You can also pass the Scalattice key on the request instead of storing it in Model Catalog: + +```python +from portkey_ai import Portkey + +portkey = Portkey( + api_key="PORTKEY_API_KEY", + provider="scalattice", + Authorization="Bearer slt_..." +) + +response = portkey.chat.completions.create( + model="qwen-3-14b", + messages=[{"role": "user", "content": "Hello!"}] +) +``` + +## Managing Scalattice Prompts + +You can manage prompts to Scalattice in the [Prompt Library](/product/prompt-library). Once a prompt is ready, use `portkey.prompts.completions.create` in your application. + +--- + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; diff --git a/product/administration/enforce-default-config.mdx b/product/administration/enforce-default-config.mdx index 39688f8c4..9dbf24284 100644 --- a/product/administration/enforce-default-config.mdx +++ b/product/administration/enforce-default-config.mdx @@ -160,6 +160,14 @@ For detailed information on API key management, refer to our API documentation: Organisation owners can set **`default_workspace_user_api_key_auto_create`** (org-only, Backend `v1.17.0+`) to control whether workspace-user API keys are auto-provisioned when users join the organisation's **default workspace**. This is separate from per-workspace **`user_api_key_auto_create`** overrides. +## Default Scopes for Auto-Created User API Keys + +Organisation and workspace admins can set **`user_api_key_auto_create_scopes`** (with workspace-level override, Backend `v1.24.2+`) to control which data-plane scopes are granted to auto-created workspace-user API keys. When set, auto-provisioned keys receive only the listed scopes instead of the full default scope set. + +- Accepts an array of valid data-plane scope strings (e.g. `["completions"]`) +- Supports workspace-level override via the `user_api_key_auto_create_scopes` workspace override toggle +- When not set, auto-created keys receive the default full scope set + ## Workspace Default Config for User API Keys Organisation admins can set a workspace-level default config for workspace-user API keys via `defaults.user_api_key_config` on workspace create/update (Backend `v1.16.2+`). diff --git a/product/enterprise-offering/components.mdx b/product/enterprise-offering/components.mdx index 721e353fd..a0fde771f 100644 --- a/product/enterprise-offering/components.mdx +++ b/product/enterprise-offering/components.mdx @@ -58,7 +58,9 @@ Portkey supports robust caching solutions to optimize performance and reduce lat - + + + diff --git a/product/enterprise-offering/org-management/directory-sync/cie-directory-sync.mdx b/product/enterprise-offering/org-management/directory-sync/cie-directory-sync.mdx new file mode 100644 index 000000000..0af124188 --- /dev/null +++ b/product/enterprise-offering/org-management/directory-sync/cie-directory-sync.mdx @@ -0,0 +1,328 @@ +--- +title: "CIE Directory Sync" +description: "Sync users and groups from Palo Alto Networks Cloud Identity Engine (CIE) into SCM's AI Gateway for automated workspace provisioning." +--- + +# CIE Directory Sync + +CIE (Cloud Identity Engine) Directory Sync allows you to pull users and groups from your organization's identity provider directories — such as **Entra ID (Azure AD)**, **Okta**, or **On-Premises Active Directory** — into SCM via Palo Alto's Cloud Identity Engine. Once synced, you can map CIE groups to SCM's AI Gateway workspaces so that users are **automatically provisioned** into the correct workspaces. + +--- + +## Overview + +CIE Directory Sync is available for organizations running in **SCM (Strata Cloud Manager)**. It replaces the need for manual user provisioning or standalone SCIM integration by leveraging CIE as the centralized identity source. + +### How It Works + +1. **CIE aggregates directories** — Your organization's identity providers (Entra ID, Okta, on-prem AD) are connected to CIE via the Strata Cloud Manager. CIE syncs and caches user/group data from these directories. +2. **Admin maps groups to workspaces** — An admin selects which CIE directory to connect, then maps CIE groups to AI Gateway workspaces. +3. **Users are auto-provisioned** — Background sync periodically pulls group membership changes from CIE and provisions/deprovisions users in the mapped workspaces automatically. + +### Key Concepts + +| Concept | Description | +|---------|-------------| +| **Domain (Connected Directory)** | An identity provider directory synced into CIE. Each domain represents a separate directory source. | +| **Tenant ID** | The CIE tenant identifier for your organization, auto-provisioned during Onboarding. You never need to enter this manually. | +| **Group** | A directory group from CIE (e.g., a security group). Groups contain users that can be mapped to workspaces. | +| **Group-Workspace Mapping** | A 1:1 link between a CIE group and an AI Gateway workspace. All members of the mapped group are automatically provisioned into that workspace. | +| **User Identity Attribute** | The CIE user attribute used as the email address — either **UPN (User Principal Name)** or **Mail (Primary Email)**. | +| **Auth Profile** | The authentication profile used to validate user identity during MCP inbound authentication. These profiles are synced from [Authentication Profiles](https://docs.paloaltonetworks.com/identity/cloud-identity-engine/authenticate-users-with-the-cloud-identity-engine) configured under CIE. | + +--- + +## Prerequisites + +Before configuring CIE Directory Sync in SCM's AI Gateway, ensure the following: + +1. **CIE is provisioned for your organization** — Your Strata Cloud Manager tenant must have CIE activated with a Directory Sync instance. This is set up during Onboarding. +2. **At least one directory is connected in CIE** — Navigate to CIE and verify that at least one directory (Entra ID, Okta, or On-Premises) has been added and has a successful sync status. +3. **You have SCM admin access** — Only organization admins can configure Directory Sync in SCM's AI Gateway. + + +CIE Directory Sync is only available for SCM Tenants. It is not available in standalone deployments. For non-SCM deployments, use [SCIM Provisioning](/product/enterprise-offering/org-management/scim/scim) instead. + + +--- + +## Setting Up Directories in CIE + +Before SCM's AI Gateway can sync from CIE, you need to connect your identity provider directories in CIE. This is done in the **Strata Cloud Manager → Cloud Identity Engine** console. + +For more information about CIE, see the [Cloud Identity Engine documentation](https://docs.paloaltonetworks.com/identity/cloud-identity-engine/cloud-identity-engine-overview). + +![CIE Directories listing — showing CIE Directory, Entra ID, and Okta directories with sync status, user/group counts, and last sync times](/images/directory-sync/cie-directories-listing.png) + +### Adding a New Directory + +1. In the CIE console, navigate to **Directory Sync → Directories**. +2. Click **Add New Directory**. +3. You will see the directory type options: + +![CIE "Set Up Directory" page — CIE Directory, On-Premises Directory, and Cloud Directory options](/images/directory-sync/cie-set-up-directory.png) + +For SCM's AI Gateway integration, the relevant directory types are: + +| Directory Type | Provider | Description | +|---------------|----------|-------------| +| **Cloud Directory** | Entra ID (Azure AD), Okta, Google | Connect a cloud identity provider. This is the most common setup. | +| **On-Premises Directory** | Active Directory | Install a Cloud Identity agent to sync from on-prem AD. | +| **CIE Directory** | CIE-native | Create a local directory managed entirely within CIE. | + + +SCM's AI Gateway can connect to **any** directory type that CIE supports. The "Connected Directory" dropdown will show all directories that have been successfully synced in CIE. + + +--- + +## Configuring Directory Sync in SCM's AI Gateway + +Navigate to **AI Gateway → Admin Settings → Authentication → Directory Sync** in the SCM console. + +![SCM AI Gateway Directory Sync configuration page — Connected Directory, User Identity Attribute, Auth Profile, Sync State, and Group Mappings](/images/directory-sync/portkey-directory-sync-overview.jpg) + +The **Configure in CIE** button redirects to your CIE Directory Sync console, where you can manage directories. + +### Step 1: Select a Connected Directory + +The **Connected Directory** dropdown shows all available directories from CIE, along with their provider type and entity counts (groups and users). + +![Connected Directory dropdown — available domains with provider type and user/group counts](/images/directory-sync/connected-directory-dropdown.png) + +Each entry displays: +- **Domain name** — the directory domain (e.g., `corp.example.com`) +- **Provider type** — `aad` (Entra ID), `okta`, `cie_directory` (CIE-native), `ad` (on-prem) +- **Group and user counts** — number of groups and users in that directory + +Select the directory you want to sync from. You can also select **None (Disconnected)** to disconnect. + + +Currently, only **one directory** can be connected at a time. + + +### Step 2: Choose the User Identity Attribute + +The **User Identity Attribute** determines which CIE attribute is used as the user's email address. + +![User Identity Attribute dropdown — UPN (User Principal Name) and Mail (Primary Email) options](/images/directory-sync/user-identity-attr-dropdown.png) + +| Attribute | Description | When to Use | +|-----------|-------------|-------------| +| **UPN (User Principal Name)** | The `userPrincipalName` attribute from the directory (e.g., `john@contoso.com`) | Default choice. Use when UPN matches the user's email. | +| **Mail (Primary Email)** | The `mail` attribute from the directory | Use when UPN differs from email (e.g., UPN is `john@contoso.onmicrosoft.com` but email is `john@contoso.com`) | + + +If the selected attribute is **missing** for a user in CIE, that user will be **skipped** during sync. Verify that your chosen attribute is populated for all users in your directory. + + + +Changing the User Identity Attribute after initial setup triggers an automatic **full re-sync** to update all user email addresses. This is safe — no users are removed during this re-sync. + + +### Step 3: Save Configuration + +Click **Save** to persist your Connected Directory and User Identity Attribute selections. + +Once saved, the sync configuration becomes **active** and the background sync scheduler begins monitoring for changes. + +### Step 4: Select an Auth Profile + +The **Auth Profile** dropdown shows authentication profiles available for your tenant. These profiles are synced from [Authentication Profiles](https://docs.paloaltonetworks.com/identity/cloud-identity-engine/authenticate-users-with-the-cloud-identity-engine) configured under CIE. The selected profile is used to validate user identity during MCP inbound authentication flows. + +--- + +## Directory Sync State + +The **Directory Sync State** section shows the current health of the sync process. + +![Directory Sync State — Status: Success, Last Updated: Sep 4, 2026, Objects Synced: 12 users · 1 groups](/images/directory-sync/sync-state-success.jpg) + +| Field | Description | +|-------|-------------| +| **Status** | Current sync status — `Success`, `In Progress`, or `Failed` | +| **Last Updated** | Timestamp of the last sync run (success or failed) | +| **Objects Synced** | Count of users and groups currently mapped to workspaces | + +### Full Sync Button + +Delta sync with CIE for group-membership updates happens every 15 minutes. So any update can take up to 15 minutes to reflect in AI Gateway. If CIE rebuilds its cache (approximately every week), AI Gateway does a full sync automatically. + +But in case there is some issue or mismatch noted and you don't want to wait for the full sync period, you can click **Full Sync** which will do a forceful full sync of data from CIE. + + +If a sync is already in progress, the full sync will run once the current sync completes. + + +--- + +## Group Mappings + +The **Group Mappings** section is where you map CIE groups to workspaces. Users in a mapped group are automatically provisioned into the corresponding workspace. + +![Group Mappings — "Default Directory" mapped to "Engineering_Workspace"](/images/directory-sync/group-mappings.jpg) + +### Adding a Mapping + +1. Click **Add Mapping**. The **Add Group Mapping** dialog opens: + +![Add Group Mapping dialog — select a CIE Group and a Workspace, then click Add](/images/directory-sync/add-group-mapping-dialog.png) + +2. Select a **CIE Group** from the dropdown. The dropdown lists all groups from your connected directory. +3. Select a **Workspace** to map the group to. +4. Click **Add**. + +### Mapping Rules + +- **1:1 mapping** — Each group can only be mapped to one workspace, and each workspace can only have one group mapped to it +- **Workspace must exist** — The target workspace must already exist. Directory Sync does not create workspaces. + +### Removing a Mapping + +Click the **delete** (trash) icon next to a mapping to remove it. Users provisioned by this mapping will be **immediately removed** from that workspace. + + +Deleting a mapping immediately removes users from the workspace. This action cannot be undone — you would need to re-create the mapping and wait for a sync to re-provision users. + + +--- + +## Viewing Directory-Provisioned Members + +Once Directory Sync is configured and group mappings are in place, users from CIE are automatically provisioned into the mapped workspaces. You can view these members through the Workspace Control page. + +### Viewing Workspaces + +Navigate to **AI Gateway → Workspace Control** to see all workspaces in your organization. + +![Workspace Control — list of all workspaces in the organization](/images/directory-sync/workspace-control-list.jpg) + +This page shows all workspaces along with their slug, creation date, and last update time. Workspaces that have CIE groups mapped to them will have directory-provisioned members automatically added. + +### Viewing Workspace Members + +Click on a workspace to open its settings, then navigate to the **Members** tab to see all members provisioned into that workspace. + +![Workspace Members — showing directory-provisioned users in Engineering_Workspace](/images/directory-sync/workspace-members.jpg) + +Each member entry shows: +- **Name** — the user's display name, derived from CIE's `Common-Name` attribute +- **Email** — the user's email, based on the User Identity Attribute you selected (UPN or Mail) +- **Created At** — when the user was provisioned into the workspace +- **Last Update** — when the user's membership was last updated by a sync + + +Members listed here are automatically managed by Directory Sync. When a user is added to or removed from the mapped CIE group, the workspace membership is updated accordingly during the next sync cycle (within 15 minutes). + + +--- + +## Creating User API Keys for Directory-Provisioned Members + +Once users are provisioned into workspaces via Directory Sync, you can create **User API Keys** scoped to individual directory-provisioned members. This allows each user to have their own key for accessing AI Gateway services within their workspace. + +### Navigating to Security Keys + +1. Select a workspace from the **Workspace** dropdown at the top of the page. +2. Navigate to **AI Gateway → Security Keys**. + +You will see the **Gateway API Keys** page with two tabs — **Service** and **User**. + +![Security Keys page — Service tab showing existing service API keys](/images/directory-sync/security-keys-service-tab.jpg) + +- **Service** keys are shared keys not tied to a specific user. +- **User** keys are tied to a specific directory-provisioned member. + +Switch to the **User** tab to view existing user API keys. + +![Security Keys — User tab showing user API keys with their owners](/images/directory-sync/security-keys-user-tab.jpg) + +### Creating a User API Key + +1. Click **+ Create New**. The **Create New Gateway API Key** form opens. + +2. Under **API Key Type**, select **User**. + +![Create API Key — Step 1: Configure API Key Details with User type selected](/images/directory-sync/create-api-key-details.jpg) + +3. Under **Select User**, choose a directory-provisioned member from the dropdown. Only users who have been synced into this workspace via Directory Sync will appear here. + +![Select User dropdown — showing directory-provisioned users](/images/directory-sync/create-api-key-select-user.png) + +4. Enter an **API Key Name** — this is required and helps identify the key later. + +5. Optionally fill in a **Short Description**, **Configuration**, and **Metadata**. + +![Filled form — User1 AIGW selected with key name "test-doc"](/images/directory-sync/create-api-key-filled.jpg) + +6. Click **Next: Set Permissions**. + +7. On the **Permissions** step, configure which permissions this key should have. Permissions are organized by resource (Agents, Completions, Logs, Mcp, Prompts) and action (Invoke, Write, Render). + +![Set up Permissions — permission matrix for the API key](/images/directory-sync/create-api-key-permissions.jpg) + +8. Click **Create Gateway API Key**. + +9. The generated API key is displayed. **Copy it now** — you will not be able to view it again. + +![Save your Gateway API Key — copy the key before closing](/images/directory-sync/create-api-key-save.png) + +10. Click **Copy and Close**. The new key will appear in the **User** tab of Security Keys. + +![Security Keys User tab — newly created key attributed to User1 AIGW](/images/directory-sync/security-keys-user-created.jpg) + + +You cannot create a User API key without selecting a user and providing a key name. Both fields are required. + + + +User API keys are scoped to the selected workspace. Each key is attributed to a specific directory-provisioned member and tracks who created it and who owns it. + + +--- + +## Disabling Directory Sync + +To disable Directory Sync entirely: + +1. Navigate to **Admin Settings → Authentication → Directory Sync** +2. Select **None (Disconnected)** from the Connected Directory dropdown + +**What happens when you disable sync:** + +1. All CIE-synced users are **removed from their mapped workspaces** +2. All group-workspace mappings are removed +3. All group and sync status records are removed +4. The sync configuration is deactivated + + +Disabling Directory Sync immediately removes all CIE-provisioned users from their workspaces. This is a destructive action. Users can be re-provisioned by re-enabling sync and re-creating group mappings. + + +To **re-enable** sync after disabling: +1. Select a Connected Directory again and save +2. Group mappings do **not** carry over — you must re-create them explicitly +3. A full sync will run to provision users into the newly mapped workspaces + +--- + +## Troubleshooting + +### Common Issues + +| Issue | Cause | Resolution | +|-------|-------|------------| +| **No domains shown** in Connected Directory dropdown | CIE not provisioned for this org, or no directories added in CIE | Ensure CIE is activated for your SCM tenant. Add directories in CIE. | +| **Users not provisioned** after mapping | Selected User Identity Attribute is missing for those users | Check CIE to confirm users have the UPN or Mail attribute populated. Switch attribute if needed. | +| **Users not removed** after deleting mapping | User belongs to another group also mapped to the same workspace | This is by design — users who have access through another mapping are not removed. | +| **Delta sync falling back to full** | CIE cache was rebuilt, or cursor expired | This is expected behavior. CIE periodically rebuilds its cache, which triggers a full resync. This acts as a self-healing mechanism. | + +--- + +## Support + +If you encounter issues with CIE Directory Sync, contact your support team. + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; + + diff --git a/product/enterprise-offering/org-management/jwt.mdx b/product/enterprise-offering/org-management/jwt.mdx index f4a3896c8..3af9d6680 100644 --- a/product/enterprise-offering/org-management/jwt.mdx +++ b/product/enterprise-offering/org-management/jwt.mdx @@ -27,7 +27,7 @@ JWT authentication can be configured under **Admin Settings** → **Organisation To validate JWTs, you must configure one of the following: -- **JWKS URL**: A URL from which the public keys will be dynamically fetched. +- **JWKS URL**: A URL from which the public keys will be dynamically fetched. You can provide multiple JWKS URLs as a comma-separated list (e.g. when rotating between IdPs); the gateway fetches and merges keys from all of them. - **JWKS JSON**: A static JSON containing public keys. ## Hard Requirements (Read First) @@ -573,8 +573,8 @@ Use this when tokens come from Okta, Auth0, Entra, Cognito, or another IdP whose | Workspace | Restrict the deployment to specific workspaces, or set an org-level default workspace. See [Workspace resolution](#workspace-resolution). | | Scopes | Register gateway scopes (e.g. `completions.write`) in your IdP, or set `JWT_LOCAL_AUTH_DEFAULT_SCOPES` on the gateway when IdP scope changes are not possible. Both `completions.write` and `portkey.completions.write` are accepted. | | Usage and rate limits | Configure limits on the workspace, or use workspace policies keyed off request metadata. | -| User attribution in logs | The gateway uses `email_id`, `sub`, or `uid` from the token for per-user logging and metrics. | -| JWKS | Point your org's JWKS URL to your IdP's discovery endpoint (e.g. `https:///.well-known/jwks.json`) in **Admin Settings** → **Organisation** → **Authentication**. | +| User attribution in logs | The gateway uses `email_id`, `sub`, or `uid` from the token for per-user logging and metrics. When the token's email matches a Portkey user in the target workspace, the request is attributed to that user (and inherits their workspace role) instead of a generic workspace service key. | +| JWKS | Point your org's JWKS URL to your IdP's discovery endpoint (e.g. `https:///.well-known/jwks.json`) in **Admin Settings** → **Organisation** → **Authentication**. Supports a comma-separated list of URLs. | ### Enable gateway-local JWT diff --git a/product/enterprise-offering/otel/complete-logs.mdx b/product/enterprise-offering/otel/complete-logs.mdx index 47dbdbec4..92712177f 100644 --- a/product/enterprise-offering/otel/complete-logs.mdx +++ b/product/enterprise-offering/otel/complete-logs.mdx @@ -124,6 +124,16 @@ Logs are exported as OpenTelemetry spans with attributes following the [GenAI se | `otel.semconv.version` | `1.40.0` | Semantic convention version | | Custom attributes | Configurable | Additional attributes from `EXPERIMENTAL_GEN_AI_OTEL_RESOURCE_ATTRIBUTES` | +### Portkey Tenant Attributes + +These attributes are set on all spans and identify the Portkey organisation, workspace, and user associated with the request: + +| Attribute | Source | Description | +|-----------|--------|-------------| +| `portkey.organisation.id` | Metrics | Organisation ID in Portkey | +| `portkey.workspace.slug` | Metrics | Slug of the workspace that made the request | +| `portkey.user.id` | Metrics | User ID or API key ID that initiated the request | + ### Common Attributes These attributes are set on all spans (inference and embeddings): diff --git a/product/mcp-gateway/authentication/cas.mdx b/product/mcp-gateway/authentication/cas.mdx new file mode 100644 index 000000000..00889d436 --- /dev/null +++ b/product/mcp-gateway/authentication/cas.mdx @@ -0,0 +1,146 @@ +--- +title: "OAuth in SCM" +sidebarTitle: "OAuth (SCM)" +description: "OAuth 2.1 authentication for MCP Gateway in SCM deployments — powered by Palo Alto Networks Cloud Authentication Service (CAS)." +--- + +CAS (Cloud Authentication Service) is the OAuth 2.1 authentication method for MCP Gateway in **SCM (Strata Cloud Manager)** deployments. When an MCP client connects to a gateway running in SCM mode, CAS handles user authentication through Palo Alto Networks' identity infrastructure instead of the AI Gateway's built-in OAuth. + + +For standard deployments, use [AI Gateway OAuth](/product/mcp-gateway/authentication/oauth) or [External OAuth](/product/mcp-gateway/authentication/external-oauth) instead. + + +--- + +## Overview + +CAS bridges the MCP Gateway's OAuth flow with Palo Alto Networks' centralized authentication. Instead of users logging in with standard credentials, they authenticate through their organization's identity provider (Entra ID, Okta, or on-prem Active Directory) via CAS. + +### How It Works + +```mermaid +sequenceDiagram + participant Client as MCP Client + participant Gateway as AI Gateway
(Data Plane) + participant CP as AI Gateway
(Control Plane) + participant CAS as CAS (Palo Alto) + participant IdP as Identity Provider + + Client->>Gateway: Connect to MCP server + Gateway->>CP: Initiate OAuth flow + CP->>CAS: Redirect user to CAS login + CAS->>IdP: Authenticate via org IdP + IdP-->>CAS: Identity verified + CAS-->>CP: User authenticated + CP-->>Gateway: Authorization code + Gateway-->>Client: Access token issued +``` + +--- + +## Prerequisites + +Before CAS authentication works for your MCP Gateway: + +1. **CIE Directory Sync configured** — Users must be provisioned into workspaces via [CIE Directory Sync](/product/enterprise-offering/org-management/directory-sync/cie-directory-sync). CAS authenticates users, but CIE is what provisions them into the system. Without CIE sync, authenticated users cannot be resolved. + +2. **Authentication Profile selected** — An Auth Profile must be selected in the CIE Directory Sync configuration. This profile determines which identity provider is used for the CAS login flow. Auth Profiles are managed in the [CIE Authentication Profiles](https://docs.paloaltonetworks.com/identity/cloud-identity-engine/authenticate-users-with-the-cloud-identity-engine) console. + + +CAS relies on CIE for user provisioning. If CIE Directory Sync is not configured, or no group-workspace mappings exist, users will authenticate successfully with CAS but fail to resolve — resulting in an authorization error. + +See [CIE Directory Sync](/product/enterprise-offering/org-management/directory-sync/cie-directory-sync) for setup instructions. + + +--- + +## Setup + +CAS authentication is automatically enabled. No additional configuration is required on the MCP Gateway side — the gateway routes authentication through CAS. + +### What the Admin Needs to Configure + +| Step | Where | What | +|------|-------|------| +| 1. Connect a directory | CIE Console | Add your identity provider (Entra ID, Okta, AD) to CIE | +| 2. Configure Directory Sync | AI Gateway → Admin Settings | Select the connected directory and user identity attribute | +| 3. Select Auth Profile | AI Gateway → Admin Settings | Choose the CAS authentication profile for MCP auth | +| 4. Map groups to workspaces | AI Gateway → Admin Settings | Map CIE groups to AI Gateway workspaces | + + +**Email claim is required.** CAS identifies users by their email address. The identity provider must include the user's email in the authentication claims. If the email claim is missing, the user cannot be resolved and authentication will fail. + +Ensure that the **User Identity Attribute** selected in CIE Directory Sync (UPN or Mail) matches the email attribute returned by your identity provider during CAS authentication. + + +--- + +## User Experience + +### First-Time Connection + +When a user connects an MCP client to the gateway for the first time, the MCP client opens a browser window and the CAS login page is shown. + +**Step 1 — Authenticate with your identity provider** + +The user enters their organization credentials on the CAS Single Sign-on page. This page is hosted by Palo Alto Networks and connects to your configured identity provider. + + +The login page appearance may vary depending on your configured Authentication Profile and identity provider. The example below shows the default CAS login for a local directory. Organizations using external IdPs (e.g., Entra ID, Okta) will see their IdP's login page instead. + + +![CAS Single Sign-on — sample login page for a local directory configuration](/images/cas-login.png) + +**Step 2 — Approve access to the MCP server** + +After authentication, the user is shown a consent page. It displays the MCP server being requested, the workspace it belongs to, and the redirect destination. The user can approve or reject the request. + + +If you are a member of multiple workspaces where the MCP server is provisioned, a workspace dropdown will appear on the consent page. Select the workspace you want the access token to be scoped to. + + +![Authorization Request — consent page showing the MCP server name, workspace selection, redirect destination, and approve/reject buttons](/images/cas-consent.png) + +Once approved, the browser redirects back to the MCP client with an access token. The client can now make MCP requests. + +### Subsequent Connections + +After the initial authentication, the MCP client uses refresh tokens to maintain access. Users are not prompted to log in again until the refresh token expires or is revoked. Approval is also remembered — subsequent connections to the same MCP server skip the consent page. + +--- + +## Comparison with Other Auth Methods + +| Feature | CAS (SCM) | AI Gateway OAuth | External OAuth | +|---------|-----------|------------------|----------------| +| **Identity provider** | Palo Alto CAS → org IdP | AI Gateway accounts | Customer's IdP | +| **User provisioning** | Via CIE Directory Sync | Self-service signup | Manual or SCIM | +| **Consent flow** | Yes | Yes | Depends on IdP | +| **Token management** | Automatic | Automatic | Customer manages | +| **MCP client support** | All standard MCP clients | All standard MCP clients | All standard MCP clients | + +--- + +## Troubleshooting + +| Issue | Cause | Resolution | +|-------|-------|------------| +| User authenticates but gets "authorization failed" | User not provisioned via CIE | Configure [CIE Directory Sync](/product/enterprise-offering/org-management/directory-sync/cie-directory-sync) and verify group-workspace mappings | +| "Email not found in claims" error | IdP not returning email attribute | Ensure the IdP includes the email claim. Verify the User Identity Attribute (UPN vs Mail) in CIE matches what the IdP returns | +| Consent page shows but redirect fails | Browser blocking the redirect | Check for browser extensions or corporate policies blocking redirects to custom URI schemes (`cursor://`, `vscode://`) | +| Token refresh fails | Refresh token expired or revoked | User must re-authenticate through the CAS flow | + +--- + +## Related + +| Topic | Description | +|-------|-------------| +| [CIE Directory Sync](/product/enterprise-offering/org-management/directory-sync/cie-directory-sync) | Provision users from CIE into AI Gateway workspaces (required for CAS) | +| [SCM Architecture](/self-hosting/hybrid-deployments/scm-architecture) | Architecture guide for SCM deployment mode | +| [OAuth](/product/mcp-gateway/authentication/oauth) | AI Gateway's built-in OAuth for non-SCM deployments | +| [External OAuth](/product/mcp-gateway/authentication/external-oauth) | Bring your own identity provider | + +import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; + + diff --git a/product/mcp-gateway/authentication/jwt.mdx b/product/mcp-gateway/authentication/jwt.mdx index f8b7de84c..e4777e3ea 100644 --- a/product/mcp-gateway/authentication/jwt.mdx +++ b/product/mcp-gateway/authentication/jwt.mdx @@ -40,6 +40,25 @@ Portkey fetches keys and caches them (default: 24 hours). When a key rotates, Po **Best for:** Most production deployments. Minimal configuration, automatic key rotation. +#### Multiple JWKS URIs + +When your infrastructure uses multiple identity providers or key sets, pass a comma-separated list of JWKS endpoints: + +```json +{ + "jwt_validation": { + "jwksUri": "https://idp-1.example.com/.well-known/jwks.json, https://idp-2.example.com/.well-known/jwks.json", + "algorithms": ["RS256"] + } +} +``` + +Portkey fetches all endpoints in parallel, merges their key sets, and finds the matching key by `kid`. If one endpoint is unavailable, tokens are still verified against the keys from the remaining endpoints. + + +URLs that contain commas in query strings (e.g., `?tenants=a,b`) are handled correctly — Portkey only splits on commas that separate distinct URLs. + + ### Inline JWKS Embed public keys directly in the configuration. Use for self-contained deployments or environments without a JWKS endpoint. diff --git a/product/mcp-gateway/guardrails.mdx b/product/mcp-gateway/guardrails.mdx index 5ddb4b65e..1bda18b8b 100644 --- a/product/mcp-gateway/guardrails.mdx +++ b/product/mcp-gateway/guardrails.mdx @@ -1,32 +1,408 @@ --- title: Guardrails -description: Apply policies to MCP requests. +description: Apply security and compliance guardrails to MCP tool calls — regex filters, content validation, webhook checks, JSON schema enforcement, and more. --- - -Guardrails for MCP Gateway is coming soon. - +When AI agents invoke tools through MCP servers, sensitive data can flow in both directions — tool inputs may contain PII or restricted content, and tool outputs may return confidential data. MCP Guardrails let you intercept and enforce policies on `tools/call` requests at two stages: -Guardrails apply policies to MCP requests before and after they execute. +| Stage | What it covers | +|-------|---------------| +| **Input** | The arguments sent _to_ the MCP tool | +| **Output** | The result returned _from_ the MCP tool | -## Available now +You can apply guardrails at the **workspace level** (default for all MCP servers in a workspace) or at the **server level** (targeting specific MCP servers and even individual tools). -**Rate limiting.** Limit requests and token consumption per API key, server, or tool. See [MCP Rate Limits](/product/mcp-gateway/rate-limits). +--- + +## How It Works + +MCP guardrails use the same policy engine as LLM guardrails, with the `target` field set to `mcp_tools`. This keeps MCP and LLM guardrails separate — existing LLM guardrails are unaffected. + +Each guardrail consists of **checks** (the validation rules) and **actions** (what happens when a check fails). You then **map** the guardrail to one or more MCP servers, choosing whether it runs on tool inputs, outputs, or both. + +When both workspace-level and server-level guardrails are configured, server-level guardrails are applied **in addition to** workspace defaults. + +--- + +## Supported Checks + +The following guardrail checks are available for MCP tool calls (`target: "mcp_tools"`): + +| Check ID | Name | Description | +|----------|------|-------------| +| `default.regexMatch` | Regex Match | Match patterns in tool inputs or outputs | +| `default.regexReplace` | Regex Replace | Match and replace patterns in tool data | +| `default.contains` | Contains | Check if content contains any, all, or none of specified words/phrases | +| `default.endsWith` | Ends With | Check if content ends with a specific string | +| `default.webhook` | Webhook | Send tool call data to an external service for custom validation | +| `default.jsonSchema` | JSON Schema | Validate tool input/output against a JSON schema | +| `default.jsonKeys` | JSON Keys | Check for required keys in JSON tool data | +| `default.sentenceCount` | Sentence Count | Enforce sentence count ranges | +| `default.wordCount` | Word Count | Enforce word count ranges | +| `default.characterCount` | Character Count | Enforce character count ranges | +| `default.containsCode` | Contains Code | Detect code (SQL, Python, TypeScript, etc.) in tool data | +| `default.validUrls` | Valid URLs | Validate that all URLs in tool data are well-formed | +| `default.isAllLowerCase` | Lowercase Check | Check if content is all lowercase | +| `default.alluppercase` | Uppercase Check | Check if content is all uppercase | +| `default.notNull` | Not Null | Ensure tool output is not null, undefined, or empty | +| `default.requiredMetadataKeys` | Required Metadata Keys | Check that metadata contains all required keys | +| `default.requiredMetadataKeyPairs` | Required Metadata Key-Value Pairs | Check for specific key-value pairs in metadata | +| `default.requestParametersCheck` | Request Parameters Check | Validate request parameters | + + +LLM-specific checks (PII detection, content moderation, language checks, and third-party provider checks like Patronus, Azure, Bedrock, etc.) are not available for MCP tool calls. Use the checks listed above for MCP guardrails. + + +--- + +## Key Concepts + +### Target + +Every guardrail has a `target` field: + +| Target | Description | +|--------|-------------| +| `llm` | Applied to LLM API requests (default, existing behavior) | +| `mcp_tools` | Applied to MCP tool calls | + +When creating a guardrail for MCP, set `target` to `"mcp_tools"`. + +### Run On + +Each MCP server mapping includes a `run_on` field that controls _when_ the guardrail executes: + +| Value | Description | +|-------|-------------| +| `input` | Run the guardrail on tool call **arguments** (before the tool executes) | +| `output` | Run the guardrail on tool call **results** (after the tool executes) | + +You can set `run_on` to `["input"]`, `["output"]`, or `["input", "output"]` for both. + +### Capability Scoping + +By default, a guardrail mapping applies to **all tools** on an MCP server. To narrow the scope to specific tools, pass `mcp_integration_capability_ids` — an array of tool capability UUIDs. The guardrail will only run on calls to those specific tools. + +--- + +## Workflow -## What's coming +### Step 1: Create an MCP Guardrail -**Pre-execution checks.** Validate tool inputs before they run. Check for sensitive data. Enforce policies before any tool executes. +Create a guardrail with `target: "mcp_tools"`. Unlike LLM guardrails, `checks` and `actions` are optional at creation time — you can configure them later. -**Content filtering.** Block requests containing sensitive data. Inspect tool outputs for PII or secrets. +```bash cURL +curl -X POST "https://api.portkey.ai/v1/guardrails" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "name": "MCP Tool Content Filter", + "target": "mcp_tools", + "checks": [ + { + "id": "default.regexMatch", + "parameters": { + "rule": "(?i)(password|secret|api_key)", + "onFail": "deny" + } + }, + { + "id": "default.contains", + "parameters": { + "words": ["DROP TABLE", "DELETE FROM"], + "operator": "none", + "onFail": "deny" + } + } + ], + "actions": { + "on_fail": "deny" + } + }' +``` -**Approval workflows.** Require explicit approval for high-risk operations like deletions or write actions. +**Response:** + +```json +{ + "id": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", + "slug": "pg-mcp-tool-content-filter-a1b2c3", + "version_id": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", + "created_at": "2025-09-03T00:00:00.000Z", + "created_by": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" +} +``` + +### Step 2: Map the Guardrail to MCP Servers + +Attach the guardrail to one or more MCP servers. You can use either the **bulk sync** endpoint (replace all mappings at once) or the **single upsert** endpoint. + +#### Bulk Sync (Recommended) + +This is a declarative, idempotent operation — it replaces the full set of MCP server mappings for the guardrail. + +```bash cURL +curl -X PUT "https://api.portkey.ai/v1/guardrails/{guardrailId}/mcp-servers" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "mcp_servers": { + "": { + "run_on": ["input", "output"] + }, + "": { + "run_on": ["input"], + "mcp_integration_capability_ids": [""] + } + } + }' +``` + +**Response:** + +```json +{ + "changed": true, + "added": 2, + "updated": 0, + "removed": 0 +} +``` + +#### Single Server Upsert + +Map or update a guardrail for a single MCP server: + +```bash cURL +curl -X PUT "https://api.portkey.ai/v1/guardrails/{guardrailId}/mcp-servers/{mcpServerId}" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "run_on": ["input", "output"], + "mcp_integration_capability_ids": ["", ""] + }' +``` + +**Response:** + +```json +{ + "map_id": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx" +} +``` + +### Step 3: Verify the Configuration + +List all MCP server mappings for a guardrail: + +```bash cURL +curl "https://api.portkey.ai/v1/guardrails/{guardrailId}/mcp-servers" \ + -H "x-portkey-api-key: $PORTKEY_API_KEY" +``` + +**Response:** + +```json +[ + { + "id": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", + "guardrail_id": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", + "mcp_server_id": "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx", + "run_on": ["input", "output"], + "capability_ids": ["xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"] + } +] +``` + +### Workspace-Level Defaults + +You can set MCP guardrails as workspace defaults so they apply to all MCP servers in a workspace. Configure `mcp_input_guardrails` and `mcp_output_guardrails` in your workspace settings. + +--- + +## Blocked Request Response + +When a guardrail blocks an MCP tool call, the gateway returns an MCP-compliant JSON-RPC error: + +```json +{ + "jsonrpc": "2.0", + "id": 1, + "error": { + "code": -32446, + "message": "Request blocked by guardrail", + "data": { + "guardrail_id": "default.regexMatch", + "reason": "Input matched restricted pattern" + } + } +} +``` + +--- -**Budget controls.** Limit usage based on cost or request volume. +## API Reference -## Get early access +### Create Guardrail + +Create a new guardrail. Set `target` to `"mcp_tools"` for MCP guardrails. + +``` +POST /v1/guardrails +``` + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `name` | string | Yes | Display name for the guardrail | +| `target` | string | No | `"llm"` (default) or `"mcp_tools"` | +| `checks` | array | Conditional | Required for `llm` target; optional for `mcp_tools` | +| `actions` | object | Conditional | Required for `llm` target; optional for `mcp_tools` | + +--- + +### List Guardrails + +Retrieve guardrails with optional filtering by target. + +``` +GET /v1/guardrails +``` + +| Parameter | Type | Required | Description | +|-----------|------|----------|-------------| +| `target` | string | No | Comma-separated: `"llm"`, `"mcp_tools"`, or both | +| `page_size` | integer | No | Results per page (default: 100) | +| `current_page` | integer | No | Page number (default: 0) | +| `search` | string | No | Search guardrails by name | +| `id` | string | No | Comma-separated guardrail IDs | + +--- + +### Get Guardrail + +Retrieve a single guardrail by ID or slug. For `mcp_tools` target guardrails, the response includes `mcp_server_mappings`. + +``` +GET /v1/guardrails/{guardrailId} +``` + +| Parameter | Type | Description | +|-----------|------|-------------| +| `guardrailId` | string | Guardrail UUID or slug (e.g., `pg-pii-filter-a1b2c3`) | + + +`mcp_server_mappings` is only included in the response when `target` is `"mcp_tools"`. + + +--- + +### Update Guardrail + +Update a guardrail's name, checks, or actions. + +``` +PUT /v1/guardrails/{guardrailId} +``` + +| Field | Type | Description | +|-------|------|-------------| +| `name` | string | Updated display name | +| `checks` | array | Updated checks configuration | +| `actions` | object | Updated actions configuration | + +--- + +### Delete Guardrail + +Archive a guardrail. This also removes all MCP server mappings. + +``` +DELETE /v1/guardrails/{guardrailId} +``` + + +A guardrail cannot be deleted if it is currently used in workspace or organisation defaults (including `mcp_input_guardrails` and `mcp_output_guardrails`). Remove it from defaults first. + + +--- + +### Bulk Sync MCP Server Mappings + +Declaratively sync all MCP server mappings for a guardrail. This replaces the entire set — servers not included in the request body are removed. + +``` +PUT /v1/guardrails/{guardrailId}/mcp-servers +``` + +**Request Body:** + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `mcp_servers` | object | Yes | Map of MCP server UUID → config | +| `mcp_servers.*.run_on` | string[] | No | `["input"]`, `["output"]`, or `["input", "output"]` (default: `["input", "output"]`) | +| `mcp_servers.*.mcp_integration_capability_ids` | string[] | No | Scope to specific tool capabilities | + +**Validation rules:** +- The guardrail must have `target: "mcp_tools"` +- All MCP server IDs must be valid UUIDs +- `run_on` must be a non-empty array containing `"input"` and/or `"output"` +- All `mcp_integration_capability_ids` must exist in the database + +--- + +### Upsert Single MCP Server Mapping + +Create or update a guardrail mapping for a single MCP server. + +``` +PUT /v1/guardrails/{guardrailId}/mcp-servers/{mcpServerId} +``` + +| Field | Type | Required | Description | +|-------|------|----------|-------------| +| `run_on` | string[] | No | `["input"]`, `["output"]`, or `["input", "output"]` (default: `["input", "output"]`) | +| `mcp_integration_capability_ids` | string[] | No | Scope to specific tool capabilities | + +--- + +### List MCP Server Mappings + +List all MCP server mappings for a guardrail. + +``` +GET /v1/guardrails/{guardrailId}/mcp-servers +``` + +--- + +## Error Responses + +| Status | Code | Description | +|--------|------|-------------| +| `400` | `validation_failed` | Invalid request body or parameters | +| `401` | `unauthorized` | Missing or invalid API key | +| `403` | `forbidden` | Insufficient permissions or feature not enabled | +| `404` | `not_found` | Guardrail, MCP server, or capability not found | +| `500` | `server_error` | Internal server error | + +--- -Contact us if you're interested in early access to MCP guardrails. +## Next Steps + + + Full list of available guardrail checks and parameters. + + + Throttle MCP tool call requests and token consumption. + + + Monitor MCP tool call logs and usage analytics. + + + Control which workspaces and users can access MCP servers. + + import PrismaAirsCta from "/snippets/prisma-airs-cta.mdx"; diff --git a/product/observability/logs-export.mdx b/product/observability/logs-export.mdx index b81665a64..905225fd1 100644 --- a/product/observability/logs-export.mdx +++ b/product/observability/logs-export.mdx @@ -458,6 +458,8 @@ For detailed API specifications, see the [Log Exports API reference](/api-refere | `POST /v1/logs/exports/{id}/start` | Start processing an export | | `GET /v1/logs/exports/{id}` | Retrieve export status | | `GET /v1/logs/exports/{id}/download` | Get signed download URL | +| `POST /v1/logs/exports/{id}/cancel` | Cancel an in-progress export | +| `DELETE /v1/logs/exports/{id}` | Delete an export job | | `GET /v1/logs/{id}` | Fetch a single log entry | ## Support