Create a chat completion

OpenAI-compatible chat completions endpoint. Accepts the OpenAI chat completions request shape and routes to any catalogued model. Also served at `/v1/openai/chat/completions`; both paths accept the same requests. This bare path matches the OpenAI SDK's default base URL shape, so pointing the SDK at `https://api-gateway.merge.dev/v1` works with no other changes.

Authentication

AuthorizationBearer
Your production key sent as a bearer token.

Request

This endpoint expects an object.
modelstringRequired
messageslist of objectsRequired
temperaturedouble or nullOptional
top_pdouble or nullOptional
ninteger or nullOptional
streambooleanOptionalDefaults to false
stopstring or list of strings or nullOptional
max_tokensinteger or nullOptional
max_completion_tokensinteger or nullOptional
presence_penaltydouble or nullOptional
frequency_penaltydouble or nullOptional
logit_biasmap from strings to doubles or nullOptional
userstring or nullOptional
toolslist of objects or nullOptional
tool_choicestring or map from strings to any or nullOptional
response_formatobject or nullOptional
OpenAI response format.
seedinteger or nullOptional
stream_optionsmap from strings to any or nullOptional
service_tierenum or nullOptional

Processing tier. ‘flex’ = discounted best-effort (slower, may be throttled); ‘priority’ = premium low-latency (not currently priced on any route, so it fails closed). Omit or ‘standard’ for normal processing. Allowed only on routes that price the requested tier, otherwise the request fails closed with 400.

Allowed values:
service_tier_fallbackbooleanOptionalDefaults to false

If the provider throttles the requested tier (429/503), retry once at the standard tier, billed at standard rates. The response service_tier shows the tier that actually served.

logprobsboolean or nullOptional
Return log probabilities of the output tokens.
top_logprobsinteger or nullOptional

Number of most likely tokens to return per position; requires logprobs.

thinkingmap from strings to any or nullOptional
Extended thinking settings for models that support it.
tagslist of objects or nullOptional

Key-value tags to attach to this request.

project_idstring or nullOptional

Project ID (UUID) to attribute this request to.

customerstring or nullOptional

Customer ID (UUID) to scope this request to, applying that customer’s routing policy, provider keys, budget, and usage attribution.

routing_policy_idstring or nullOptional

Routing policy ID for this request: an org-level policy, or a project or customer policy when the request carries that scope.

priority_orderlist of strings or objects or nullOptional

Ordered list of models to try, first success wins, with the same rules as priority_order on POST /v1/responses.

vendorstring or nullOptional

Restrict routing to one vendor (hosting platform).

vendorslist of strings or nullOptional

Ordered list of acceptable vendors; the first available one wins.

provider_optionsmap from strings to any or nullOptional

Provider-specific options, keyed by provider family.

prompt_cache_keystring or nullOptional

Prompt-cache routing hint forwarded to providers that cache automatically.

include_routing_metadatabooleanOptionalDefaults to false
Return routing metadata for this request.
include_guardrails_metadatabooleanOptionalDefaults to false

Return prompt-injection and DLP results in a top-level guardrails object.

Response

Successful Response

Errors

422
Unprocessable Entity Error