For AI agents: a documentation index is available at the root level at /llms.txt. Append /llms.txt to any URL for a page-level index, or .md for the markdown version of any page.
OpenAI-compatible chat completions endpoint. Accepts the OpenAI chat completions request shape and routes to any catalogued model.
Also served at `/v1/openai/chat/completions`; both paths accept the same requests. This bare path matches the OpenAI SDK's default base URL shape, so pointing the SDK at `https://api-gateway.merge.dev/v1` works with no other changes.
Authentication
AuthorizationBearer
Your production key sent as a bearer token.
Request
This endpoint expects an object.
modelstringRequired
messageslist of objectsRequired
temperaturedouble or nullOptional
top_pdouble or nullOptional
ninteger or nullOptional
streambooleanOptionalDefaults to false
stopstring or list of strings or nullOptional
max_tokensinteger or nullOptional
max_completion_tokensinteger or nullOptional
presence_penaltydouble or nullOptional
frequency_penaltydouble or nullOptional
logit_biasmap from strings to doubles or nullOptional
userstring or nullOptional
toolslist of objects or nullOptional
tool_choicestring or map from strings to any or nullOptional
response_formatobject or nullOptional
OpenAI response format.
seedinteger or nullOptional
stream_optionsmap from strings to any or nullOptional
service_tierenum or nullOptional
Processing tier. ‘flex’ = discounted best-effort (slower, may be throttled); ‘priority’ = premium low-latency (not currently priced on any route, so it fails closed). Omit or ‘standard’ for normal processing. Allowed only on routes that price the requested tier, otherwise the request fails closed with 400.
Allowed values:
service_tier_fallbackbooleanOptionalDefaults to false
If the provider throttles the requested tier (429/503), retry once at the standard tier, billed at standard rates. The response service_tier shows the tier that actually served.
logprobsboolean or nullOptional
Return log probabilities of the output tokens.
top_logprobsinteger or nullOptional
Number of most likely tokens to return per position; requires logprobs.
thinkingmap from strings to any or nullOptional
Extended thinking settings for models that support it.
tagslist of objects or nullOptional
Key-value tags to attach to this request.
project_idstring or nullOptional
Project ID (UUID) to attribute this request to.
customerstring or nullOptional
Customer ID (UUID) to scope this request to, applying that customer’s routing policy, provider keys, budget, and usage attribution.
routing_policy_idstring or nullOptional
Routing policy ID for this request: an org-level policy, or a project or customer policy when the request carries that scope.
priority_orderlist of strings or objects or nullOptional
Ordered list of models to try, first success wins, with the same rules as priority_order on POST /v1/responses.
vendorstring or nullOptional
Restrict routing to one vendor (hosting platform).
vendorslist of strings or nullOptional
Ordered list of acceptable vendors; the first available one wins.
provider_optionsmap from strings to any or nullOptional
Provider-specific options, keyed by provider family.
prompt_cache_keystring or nullOptional
Prompt-cache routing hint forwarded to providers that cache automatically.
include_routing_metadatabooleanOptionalDefaults to false
Return routing metadata for this request.
include_guardrails_metadatabooleanOptionalDefaults to false
Return prompt-injection and DLP results in a top-level guardrails object.
Response
Successful Response
Errors
422
Unprocessable Entity Error
OpenAI-compatible chat completions endpoint. Accepts the OpenAI chat completions request shape and routes to any catalogued model.
Also served at /v1/openai/chat/completions; both paths accept the same requests. This bare path matches the OpenAI SDK’s default base URL shape, so pointing the SDK at https://api-gateway.merge.dev/v1 works with no other changes.