API details

Overview of the Merge Gateway API

Base API URL

All API endpoints in the reference documentation are relative to the following base URL:

https://api-gateway.merge.dev/v1

Authentication

For any request you make when communicating with Merge Gateway, you will need an API key to authenticate yourself as an authorized user.

Add your API Key with a “Bearer ” prefix as a header called Authorization to authorize your Merge API requests. This header must be included in every request in this format:

Authorization: Bearer <your_api_key>

You do not have to use the dashboard to get a key. The Management API mints, rotates, and revokes API keys programmatically, which is what you want if keys are provisioned per customer or checked into infrastructure-as-code.

Key endpoints

Gateway’s model-calling surface is centered around these endpoint groups:

  • GET /models
    • list models
    • filter by provider or vendor
    • fetch a single model by query string with ?model=<provider/model_id>
  • GET /vendors
    • list execution vendors and the models they currently serve
    • fetch a single vendor with GET /vendors/{vendor_id}
  • POST /responses
    • create an LLM response
    • the response includes the vendor that ultimately served the request and the service_tier that actually served it
  • POST /embeddings
    • create embeddings, priced on input tokens

Gateway also exposes SDK-compatible surfaces under /v1/openai, /v1/anthropic, /v1/ai-sdk, and /v1/langchain, so an existing client reaches the same models with a base URL change. See Get started.

Administering the org, rather than calling models, is a separate surface: see Management API below.

Management API

API keys, projects, routing policies, and usage are administered programmatically through the Management API, on the same host. It is authenticated by a management key (Authorization: Bearer mgmt_<your_management_key>) rather than a gateway API key, so the credential that provisions keys is never one that can make model calls. Create a management key in the dashboard under API keys → Management keys.

Endpoint groupWhat it covers
/v1/keysCreate, list, update, and revoke API keys, each with an optional spend limit and a limit_reset window of daily, weekly, or monthly
/v1/projectsCreate and manage projects, their routing policy, budgets, usage, and per-project prompt injection and DLP settings
/v1/routing-policiesCreate and update routing policies and set the org default
/v1/organization/usageOrg-wide and per-project spend rollups with per-model breakdowns

Full reference and scopes: Management API. Walkthroughs: API keys and Projects API.

Models API shape

GET /models returns canonical model identity at the top level and vendor-specific execution metadata under vendors.

Example:

1{
2 "model": "anthropic/claude-opus-4-6",
3 "provider": "anthropic",
4 "display_name": "Claude Opus 4.6",
5 "vendors": {
6 "anthropic": {
7 "launch_date": "2025-05-14",
8 "context_window": 1000000,
9 "max_output_tokens": 32768,
10 "availability_status": "available",
11 "capabilities": {
12 "input": ["text", "image"],
13 "output": ["text", "tool_use"],
14 "supports_tool_calling": true,
15 "supports_tool_choice": true,
16 "supports_structured_outputs": true,
17 "streaming": true
18 },
19 "service_tiers": ["standard", "flex"],
20 "pricing": {
21 "input_per_million": 2.5,
22 "output_per_million": 10,
23 "currency": "USD",
24 "flex": { "input_per_million": 1.25, "output_per_million": 7.5 }
25 }
26 }
27 },
28 "availability_status": "available",
29 "created_at": "2025-05-14T00:00:00Z",
30 "updated_at": "2026-03-01T00:00:00Z"
31}

Use GET /models?model=<provider/model_id> when you want one specific model object but prefer a query parameter over a path parameter. The slash is fine there because it is part of the query parameter value.

Vendors API shape

GET /vendors returns execution hosts, not canonical model owners.

Example:

1{
2 "vendor": "bedrock",
3 "name": "AWS Bedrock",
4 "models": [
5 "anthropic/claude-opus-4-6",
6 "google/gemma-3-27b-it"
7 ],
8 "supports_zdr": true,
9 "supports_byok": true,
10 "availability_status": "active"
11}

Pagination

GET /models and GET /vendors return a paged list envelope:

1{
2 "object": "list",
3 "data": [ ... ],
4 "has_more": true,
5 "next_cursor": "mistral/magistral-medium-2509-thinking"
6}

limit sets the page size. It defaults to 50 and caps at 500. A higher value is rejected before the request reaches the catalog:

1HTTP 422
2{
3 "detail": [
4 {
5 "type": "less_than_equal",
6 "loc": ["query", "limit"],
7 "msg": "Input should be less than or equal to 500",
8 "input": "1000",
9 "ctx": { "le": 500 }
10 }
11 ]
12}

Do not size a single request to the catalog. The catalog grows as models are added, so a limit that fits today silently starts truncating later. Page instead: send the previous response’s next_cursor back as cursor and keep going while has_more is true. The next page starts after the cursor entry, so pages never overlap, and next_cursor is null on the last page. Treat the cursor as opaque, a value to pass back rather than one to construct.

Python
1import requests
2
3BASE = "https://api-gateway.merge.dev/v1"
4headers = {"Authorization": "Bearer YOUR_API_KEY"}
5
6models, cursor = [], None
7while True:
8 params = {"limit": 500}
9 if cursor:
10 params["cursor"] = cursor
11 page = requests.get(f"{BASE}/models", headers=headers, params=params).json()
12 models.extend(page["data"])
13 if not page["has_more"]:
14 break
15 cursor = page["next_cursor"]
16
17print(len(models))
cURL
$# First page
$curl "https://api-gateway.merge.dev/v1/models?limit=500" \
> -H "Authorization: Bearer YOUR_API_KEY"
$
$# Next page: pass the previous response's next_cursor
>curl "https://api-gateway.merge.dev/v1/models?limit=500&cursor=mistral/magistral-medium-2509-thinking" \
> -H "Authorization: Bearer YOUR_API_KEY"

Filters do not carry across pages on their own, so repeat provider or vendor on every request in the loop alongside cursor. Fetching one model with ?model=<provider/model_id> returns a bare model object rather than a list envelope, so it has no has_more or next_cursor.

The Management API paginates separately and caps its list endpoints at 100 per page: GET /v1/keys takes offset and limit and reports has_more, and GET /v1/projects takes cursor and limit.

Responses

POST /responses returns the canonical model that served the request and a top-level vendor field for the execution host that actually handled it.

Request parameters

Beyond the message input, the request accepts these optional fields:

FieldTypeDescription
service_tier"standard" | "flex" | "priority"Processing tier. Omit for standard. Only allowed on routes priced for the tier, otherwise the request fails closed with 400. priority is accepted but not currently priced on any route, so it always fails closed today.
service_tier_fallbackbooleanIf the provider throttles the requested tier (429 / 503), retry once at standard instead of erroring. Defaults to false.

Served tier and billing

The response echoes a top-level service_tier: the tier that actually served the request and the rate you were billed at. On a throttle fallback this differs from the tier you sent (e.g. you requested flex but were served and billed at standard).

1{
2 "model": "openai/gpt-5.4",
3 "vendor": "openai",
4 "service_tier": "flex",
5 "usage": { "input_tokens": 812, "output_tokens": 180, "total_tokens": 992, "cost": 0.00219 }
6}

See Service tiers for the full flow, supported vendor routes, and fallback behavior.

Per-call cost

Every response carries usage.cost: the provider cost in USD that Gateway computed for that call, from the tokens above and the pricing of the route that actually served the request. No request parameter is needed.

1"usage": { "input_tokens": 812, "output_tokens": 180, "total_tokens": 992, "cost": 0.00219 }

usage.cost is returned on:

EndpointNotes
POST /responsesStreaming responses carry usage only on the final response.done chunk, so that is where cost appears
POST /embeddingsPriced on input tokens
POST /openai/responses, POST /openai/chat/completions, POST /openai/embeddingsSame value on the OpenAI-compatible surfaces, including the final streaming chunk

Three details worth knowing:

  • cost is null, never 0, when the served route has no pricing on file. A false zero would be recorded as real spend by anything summing costs, so an unknown price is reported as unknown. Cost-tracking tools that read this field, such as Langfuse, skip a null instead of logging $0.
  • It is the same figure as routing.cost_usd, which is only returned when you set include_routing_metadata: true. Both come from the same calculation, so a request that asks for both sees them agree. Use usage.cost for everyday cost tracking and routing metadata when you also need the policy and vendor decisions behind the call.
  • It is provider cost, not your invoiced amount. The two can differ, and the invoice is the authority. Billing starts from this figure and then applies your organization’s plan rate, prices BYOK requests against the model’s list price rather than what your own provider account charged, and adds server-tool charges such as web search, which never appear in usage.cost.

Cost reflects the tier that served the request, so a flex request that fell back to standard is priced at standard. For invoiced spend, budgets, and dashboards, see Cost governance and savings.