Create response

Create an LLM response in Gateway's native Responses shape. This is not OpenAI's Responses API: clients that implement it (the OpenAI SDK's `client.responses.*`, Codex) should use `POST /v1/openai/responses`, or send `X-Merge-Wire-Format: openai` on this endpoint to receive the OpenAI shape. Codex is detected by its `originator` header and served the OpenAI shape automatically. OpenAI-only request fields with a native equivalent are translated (`instructions` → leading system message, `reasoning.effort` → reasoning effort, `text.format` → `response_format`, `max_output_tokens` → `max_tokens`); any other OpenAI-only field is ignored and reported in `warnings` with code `openai_fields_ignored`.

Authentication

AuthorizationBearer
Your production key sent as a bearer token.

Request

This endpoint expects an object.
inputlist of objectsRequired
Conversation history
modelstring or nullOptional

Model ID as provider/model-name or a bare name, an @alias/<slug> model alias, or default_routing; optional when a routing policy applies.

customerstring or nullOptional

Customer ID (UUID) to scope this request to, applying that customer’s routing policy, provider keys, budget, and usage attribution. A non-UUID value returns 422 and an unknown customer returns 404 customer_not_found.

routing_policy_idstring or nullOptional

Routing policy ID for this request: an org-level policy, or a project or customer policy when the request carries that scope.

priority_orderlist of strings or objects or nullOptional

An inline PRIORITY list: the models to try for this request in order, first success wins. Each entry is a model id (provider/model or a bare name) or a {model, priority, vendors} object. Up to 10 entries. Bypasses every stored routing policy (org default, customer default, key binding) while the key’s allowed models, vendor and region limits, ZDR and vendor access still apply per entry. Fails over on provider faults only (5xx, 429, timeouts, 402); on a stream, before the first byte. The list is materialized as a routing policy with an id derived from its content and scope, attributed in routing.policy_id and the x-merge-routing-policy-id header, and reusable later via routing_policy_id. model must be omitted or equal the first entry. Cannot be combined with routing_policy_id, a policy alias, or the top-level vendor / vendors pins. Also accepted on every SDK-compatible surface. See Sending a priority list with the request.

toolslist of objects or nullOptional

Tools the model may call: function tools, or provider-native tools passed through by type.

tool_choicestring or map from strings to any or nullOptional

Tool choice policy: ‘auto’, ‘none’, ‘required’, or {‘type’: ‘function’, ‘function’: {‘name’: ’…’}}

max_tokensinteger or nullOptional
Maximum tokens to generate
temperaturedouble or nullOptional0-2
Sampling temperature
top_pdouble or nullOptional0-1
Nucleus sampling parameter
stoplist of strings or nullOptional
Stop sequences
response_formatobject or nullOptional

Response format (json_object or json_schema structured output)

streambooleanOptionalDefaults to false
Whether to stream the response
tagslist of objects or nullOptional

Tags to attach to this request for categorization (key-value pairs)

project_idstring or nullOptional

Optional project ID (UUID) to associate with this request

include_routing_metadatabooleanOptionalDefaults to false
Include detailed routing metadata in response
include_guardrails_metadatabooleanOptionalDefaults to false

Include prompt-injection protection and DLP results for this request in a top-level guardrails object. In prompt-injection alert mode this also runs the direct scan on the request path (bounded by the alert-mode budget) so the score is available in the body; requests without the flag are unaffected.

vendorstring or nullOptional

Restrict routing to a specific vendor (hosting platform)

vendorslist of strings or nullOptional
Ordered list of acceptable vendors. First available wins.
service_tierenum or nullOptional

Processing tier: standard (default), flex, priority (fast is an alias), or ultrafast (OpenAI only), where default, auto, unspecified, and standard_only also mean standard; a tier the route doesn’t price returns 400 unsupported_params.

service_tier_fallbackbooleanOptionalDefaults to false

If the provider throttles the requested tier (429/503), retry once at the standard tier, billed at standard rates. The response service_tier shows the tier that actually served.

previous_response_idstring or nullOptional

ID of a stored response (resp_...) whose conversation is replayed before this request’s input; an expired or unknown ID returns 404 response_not_found.

storeboolean or nullOptional

Set true to store this turn so a later request can continue it with previous_response_id; stored turns expire an hour after their last use and are never kept under zero data retention.

logprobsboolean or nullOptional
Return log probabilities of the output tokens.
top_logprobsinteger or nullOptional

Number of most likely tokens to return per position; requires logprobs.

userstring or nullOptional

End-user identifier forwarded to the provider.

safety_identifierstring or nullOptional

End-user identifier forwarded to providers that support it (OpenAI’s newer name for user).

prompt_cache_retentionstring or nullOptional

Prompt-cache retention hint forwarded to OpenAI.

parallel_tool_callsboolean or nullOptional

Whether the model may make several tool calls in one turn; omit to keep the provider default.

modalitieslist of enums or nullOptional

Output modalities to generate; include image for inline image generation on a route that supports image output, otherwise the request returns 400.

Allowed values:
cache_controlmap from strings to any or nullOptional

Top-level prompt-caching setting for providers that cache automatically.

session_idstring or nullOptional

Caching session hint, up to 256 characters, also accepted as the X-Session-Id header, which wins when both are set.

thinkingobject or nullOptional

Extended thinking settings; the model can reason between tool calls.

provider_optionsmap from strings to any or nullOptional

Provider-specific options, keyed by provider family.

fusionobject or nullOptional

Multi-model fusion settings, used when model is fusion.

Response

Successful Response
idstring
created_atdatetime
modelstring

The model that generated the response, in ‘provider/model-name’ format

outputlist of objects
usageobject
Token usage statistics.
object"response"OptionalDefaults to response
vendorstring or nullOptional
The execution vendor that served the request
provider_request_idstring or nullOptional
The upstream provider's request ID
routingobject or nullOptional

Routing metadata (only present when include_routing_metadata=true)

guardrailsobject or nullOptional

Prompt-injection protection and DLP results (only present when include_guardrails_metadata=true). Also present on blocked 422 responses beside error.

service_tierstring or nullOptional
The tier that actually served the request and was billed. Differs from the requested tier on a throttle fallback.

Errors

422
Unprocessable Entity Error