API details
This page covers what every Gateway endpoint shares. The endpoint reference in this tab has the full schemas.
Base URL and authentication
Send a Gateway API key from the API keys page as a bearer token:
The Management API administers keys, projects, routing policies, and usage with a separate mgmt_ key on the same host. There’s no rotate endpoint: create a new key, then delete or disable the old one.
Endpoints
SDK-compatible surfaces under /v1/openai, /v1/anthropic, /v1/ai-sdk, and /v1/langchain need only a base URL change in an existing client (Get started).
/v1/responses is Gateway's native API, not OpenAI's Responses API
The native endpoint streams response.stream and response.done snapshots; OpenAI’s streams response.created through response.completed. Point OpenAI Responses clients, including the OpenAI SDK’s client.responses.* and Codex, at /v1/openai/responses (details).
Models API shape
GET /v1/models returns canonical model identity at the top level and per-vendor execution details under vendors:
Models that need provider access enabled are still listed, with access_required: true and access_reason: "vendor_access_required".
The listing shows only what the caller can route to, applying your organization’s vendor and region controls, blocklist, and BYOK-only setting, the key’s project and allowed_models, zero data retention (non-ZDR vendors and models without a ZDR route are omitted), and a customer’s controls when the request carries a customer API key or X-Merge-Customer. An exhausted customer budget doesn’t change it.
GET /v1/vendors returns execution hosts, such as bedrock (“Amazon Bedrock”), each with models, supports_zdr, supports_byok, and availability_status (active, degraded, or unavailable).
Pagination
GET /v1/models and GET /v1/vendors return {"object": "list", "data": [...], "has_more", "next_cursor"}. limit defaults to 50 and caps at 500 (larger returns 422). The catalog grows, so page with cursor: pass the previous next_cursor while has_more is true, repeating any filters. Pages never overlap, but an unknown or stale cursor restarts at the first page, so stop if you see a model ID twice.
The Management API paginates separately, up to 100 per page: GET /v1/keys uses offset and limit, and GET /v1/projects uses cursor and limit.
Shared request fields
All work on POST /v1/responses; customer, routing_policy_id, priority_order, and the two metadata flags also work on every SDK-compatible surface.
Sampling on fixed-sampling models. Gateway strips sampling fields a model would reject, so it runs at its own defaults and temperature: 0 isn’t deterministic.
When a vendor rejects stop sequences, stop is removed and the request retried. top_k is honored only on the Anthropic surface.
Headers
X-Customer-ID is a free-form tracing tag and doesn’t scope the request.
OpenAI Responses API clients
OpenAI’s Responses API is at POST /v1/openai/responses. For clients that reach native POST /v1/responses by mistake:
- Codex requests (an
originatorheader starting withcodex) get the OpenAI shape automatically X-Merge-Wire-Format: openaireturns the OpenAI shape for any caller. Any other value pins the native shape, overriding Codex detection.- OpenAI-only fields with a native equivalent are translated:
instructionsbecomes a leading system message,reasoning.effortandreasoning.summaryset reasoning,text.formatbecomesresponse_format, andmax_output_tokensis an alias ofmax_tokens. A native field wins over its OpenAI counterpart. - Other OpenAI-only fields (
background,include,max_tool_calls,metadata,prompt,prompt_cache_key,stream_options,truncation) are ignored and listed in anopenai_fields_ignoredwarning
An unknown reasoning.effort or a malformed text.format returns 400.
OpenAI hosted tools on POST /v1/openai/responses
OpenAI’s hosted tools run only for an openai/ model on the direct openai vendor:
Anthropic’s tool search tools pass through on /v1/messages, including Bedrock-hosted Claude.
Per-call cost
usage.cost is the call’s USD cost at the pricing of the route and tier that served it, including any model discount or committed-spend contract and OpenAI hosted image generation. It’s on /v1/responses, embeddings, and the OpenAI, AI SDK, and LangChain surfaces (not Anthropic), and arrives on a stream’s terminal frame.
The top-level service_tier is the tier that served and billed the request, so a flex request that fell back to standard reports and costs standard.