API details

Overview of the Merge Gateway API

This page covers what every Gateway endpoint shares. The endpoint reference in this tab has the full schemas.

Base URL and authentication

https://api-gateway.merge.dev/v1

Send a Gateway API key from the API keys page as a bearer token:

Authorization: Bearer mg_...

The Management API administers keys, projects, routing policies, and usage with a separate mgmt_ key on the same host. There’s no rotate endpoint: create a new key, then delete or disable the old one.

Endpoints

EndpointPurpose
GET /v1/models, GET /v1/models/{model_id}The model catalog, filterable by provider, vendor, and zero_data_retention. Pass ?model=provider/model to retrieve one model.
GET /v1/models/discountedModels with a current discount
GET /v1/vendors, GET /v1/vendors/{vendor_id}Execution vendors and the models they serve, filterable by zero_data_retention
POST /v1/responsesGateway’s native Responses API
GET /v1/responses/{response_id}, DELETE /v1/responses/{response_id}Read or delete a response stored with store: true. Stored responses expire an hour after last use, and previous_response_id resumes one.
POST /v1/chat/completionsOpenAI Chat Completions shape, also served at /v1/openai/chat/completions
POST /v1/messages, POST /v1/messages/count_tokensAnthropic Messages shape, also served at /v1/anthropic/v1/messages
POST /v1/embeddingsEmbeddings, priced on input tokens
POST /v1/decisionsTyped answers from a decision model
POST /v1/images/generations, POST /v1/images/editsImage generation and editing
POST /v1/audio/speech, POST /v1/audio/transcriptionsSpeech synthesis and transcription
POST /v1/videos, GET /v1/videos, GET /v1/videos/{video_id}, DELETE /v1/videos/{video_id}, GET /v1/videos/{video_id}/contentAsync video generation
/v1/batchesBatch inference at the batch tier
GET /v1/tagsThe tag key-value pairs available to your organization
GET /v1/routing/policies, GET /v1/routing/strategiesYour routing policies and the available strategies

SDK-compatible surfaces under /v1/openai, /v1/anthropic, /v1/ai-sdk, and /v1/langchain need only a base URL change in an existing client (Get started).

/v1/responses is Gateway's native API, not OpenAI's Responses API

The native endpoint streams response.stream and response.done snapshots; OpenAI’s streams response.created through response.completed. Point OpenAI Responses clients, including the OpenAI SDK’s client.responses.* and Codex, at /v1/openai/responses (details).

Models API shape

GET /v1/models returns canonical model identity at the top level and per-vendor execution details under vendors:

{
"model": "openai/gpt-5.4",
"provider": "openai",
"display_name": "GPT-5.4",
"aliases": [],
"vendors": {
"openai": {
"launch_date": "2026-03-05",
"context_window": 1050000,
"max_output_tokens": 128000,
"availability_status": "available",
"zero_data_retention": true,
"capabilities": {
"input": ["text", "image", "document"],
"output": ["text", "tool_use"],
"supports_tool_calling": true,
"supports_structured_outputs": true,
"supports_reasoning": true,
"reasoning": { "effort_values": ["none", "low", "medium", "high", "xhigh"] },
"streaming": true
},
"service_tiers": ["standard", "flex", "priority", "batch"],
"pricing": {
"currency": "USD",
"input_per_million": 2.5,
"output_per_million": 15,
"cache_read_per_million": 0.25,
"flex": { "input_per_million": 1.25, "output_per_million": 7.5 },
"priority": { "input_per_million": 5, "output_per_million": 30 },
"batch": { "input_per_million": 1.25, "output_per_million": 7.5 }
},
"batch_endpoints": ["/v1/chat/completions", "/v1/embeddings", "/v1/responses"]
}
},
"availability_status": "available",
"access_required": false,
"access_reason": null
}

Models that need provider access enabled are still listed, with access_required: true and access_reason: "vendor_access_required".

The listing shows only what the caller can route to, applying your organization’s vendor and region controls, blocklist, and BYOK-only setting, the key’s project and allowed_models, zero data retention (non-ZDR vendors and models without a ZDR route are omitted), and a customer’s controls when the request carries a customer API key or X-Merge-Customer. An exhausted customer budget doesn’t change it.

GET /v1/vendors returns execution hosts, such as bedrock (“Amazon Bedrock”), each with models, supports_zdr, supports_byok, and availability_status (active, degraded, or unavailable).

Pagination

GET /v1/models and GET /v1/vendors return {"object": "list", "data": [...], "has_more", "next_cursor"}. limit defaults to 50 and caps at 500 (larger returns 422). The catalog grows, so page with cursor: pass the previous next_cursor while has_more is true, repeating any filters. Pages never overlap, but an unknown or stale cursor restarts at the first page, so stop if you see a model ID twice.

curl "https://api-gateway.merge.dev/v1/models?limit=500" \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY"
# Next page: pass the previous response's next_cursor
curl "https://api-gateway.merge.dev/v1/models?limit=500&cursor=mistral/magistral-medium-2509-thinking" \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY"

The Management API paginates separately, up to 100 per page: GET /v1/keys uses offset and limit, and GET /v1/projects uses cursor and limit.

Shared request fields

All work on POST /v1/responses; customer, routing_policy_id, priority_order, and the two metadata flags also work on every SDK-compatible surface.

FieldDescription
service_tierstandard (default), flex, priority, or ultrafast (direct OpenAI routes). fast means priority, and default, auto, unspecified, and standard_only mean standard. A tier the route doesn’t price returns 400 unsupported_params. batch applies only to /v1/batches. See Service tiers.
service_tier_fallbacktrue retries once at standard when the vendor throttles the tier with a 429 or 503. On streams, only before the first token. Defaults to false.
customerEmbedded Routing Stack customer UUID. Its routing policies, vendor keys, budget, and usage apply
routing_policy_idA routing policy UUID to use for this request
priority_orderAn inline priority list of up to 10 models, each a model ID or {model, priority, vendors}. It can’t be combined with routing_policy_id, vendor, or vendors. See Sending a priority list with the request.
include_routing_metadataAdds a routing object with the policy, vendor, and cost decisions
include_guardrails_metadataAdds a guardrails object with the prompt injection and DLP outcome, including on a blocked 422. It never contains prompt text.
logprobs, top_logprobsToken log probabilities in OpenAI’s shape. A route that refuses them is served without them.
parallel_tool_callsfalse limits the model to one tool call per turn

Sampling on fixed-sampling models. Gateway strips sampling fields a model would reject, so it runs at its own defaults and temperature: 0 isn’t deterministic.

ModelsDropped
Newest Claude modelstemperature, top_p, and top_k, always
OpenAI GPT-5.4 and later reasoning modelstemperature and top_p unless 1
Kimi K2.5 and K3 on moonshottemperature unless 1

When a vendor rejects stop sequences, stop is removed and the request retried. top_k is honored only on the Anthropic surface.

Headers

HeaderDirectionMeaning
x-merge-vendor, x-merge-modelResponseThe vendor and canonical model that served a non-streaming request
x-merge-routing-policy-idResponseThe policy that resolved, on non-streaming requests. On streams, read routing from the terminal frame.
X-Request-IDResponseGateway’s request ID, for support
X-Merge-Customer, X-Merge-Routing-Policy-IdRequestHeader forms of customer and routing_policy_id. A header that conflicts with the body or a customer API key returns 400 customer_mismatch or routing_policy_mismatch, and a non-UUID returns 400 invalid_customer or invalid_routing_policy_id. The routing policy header applies to chat endpoints only.
include-routing-metadata, X-Merge-Include-Routing, include-guardrails-metadataRequestTurn metadata on (value true). Headers only turn it on, never off. Underscore spellings also work, but nginx drops them by default.
X-Merge-Wire-FormatRequestopenai or native on POST /v1/responses. See below.

X-Customer-ID is a free-form tracing tag and doesn’t scope the request.

OpenAI Responses API clients

OpenAI’s Responses API is at POST /v1/openai/responses. For clients that reach native POST /v1/responses by mistake:

  • Codex requests (an originator header starting with codex) get the OpenAI shape automatically
  • X-Merge-Wire-Format: openai returns the OpenAI shape for any caller. Any other value pins the native shape, overriding Codex detection.
  • OpenAI-only fields with a native equivalent are translated: instructions becomes a leading system message, reasoning.effort and reasoning.summary set reasoning, text.format becomes response_format, and max_output_tokens is an alias of max_tokens. A native field wins over its OpenAI counterpart.
  • Other OpenAI-only fields (background, include, max_tool_calls, metadata, prompt, prompt_cache_key, stream_options, truncation) are ignored and listed in an openai_fields_ignored warning

An unknown reasoning.effort or a malformed text.format returns 400.

OpenAI hosted tools on POST /v1/openai/responses

OpenAI’s hosted tools run only for an openai/ model on the direct openai vendor:

ToolBehavior
web_search, web_search_previewStreaming or not. Elsewhere removed with an openai_web_search_dropped warning; use {"type": "merge:web_search"} for search on any model.
image_generationRemoved elsewhere with an openai_image_generation_dropped warning
tool_searchServer execution runs on OpenAI, and its items round-trip on the next turn. A request that can’t reach OpenAI returns 400 openai_tool_search_requires_openai. With "execution": "client", the protocol Codex uses, it works on every model.
code_interpreter, file_search, mcp, local_shellNever run. /v1/openai/responses removes them with a tools_dropped warning, and native /v1/responses returns 400 unsupported_tool_type.

Anthropic’s tool search tools pass through on /v1/messages, including Bedrock-hosted Claude.

Per-call cost

usage.cost is the call’s USD cost at the pricing of the route and tier that served it, including any model discount or committed-spend contract and OpenAI hosted image generation. It’s on /v1/responses, embeddings, and the OpenAI, AI SDK, and LangChain surfaces (not Anthropic), and arrives on a stream’s terminal frame.

{
"model": "openai/gpt-5.4",
"vendor": "openai",
"service_tier": "flex",
"usage": { "input_tokens": 812, "output_tokens": 180, "total_tokens": 992, "cost": 0.002365 }
}

The top-level service_tier is the tier that served and billed the request, so a flex request that fell back to standard reports and costs standard.

DetailBehavior
Unpriced routecost is null, never 0
routing.cost_usdThe same figure, returned with include_routing_metadata
routing.merge_fee_usdMerge’s fee at your plan rate, or 0 when no markup applies, returned with include_routing_metadata. On BYOK it’s computed on the model’s list price, not on cost_usd. null when cost_usd is null.
Server toolsWeb search and other server-tool charges aren’t in cost or merge_fee_usd
InvoiceThe invoice is the authority. It applies your plan rate, prices BYOK requests at list price, and adds server-tool charges. Logs shows both halves per request. See Cost governance.

Next steps