Create response
Authentication
Request
Model ID as provider/model-name or a bare name, an @alias/<slug> model alias, or default_routing; optional when a routing policy applies.
Customer ID (UUID) to scope this request to, applying that customer’s routing policy, provider keys, budget, and usage attribution. A non-UUID value returns 422 and an unknown customer returns 404 customer_not_found.
Routing policy ID for this request: an org-level policy, or a project or customer policy when the request carries that scope.
An inline PRIORITY list: the models to try for this request in order, first success wins. Each entry is a model id (provider/model or a bare name) or a {model, priority, vendors} object. Up to 10 entries. Bypasses every stored routing policy (org default, customer default, key binding) while the key’s allowed models, vendor and region limits, ZDR and vendor access still apply per entry. Fails over on provider faults only (5xx, 429, timeouts, 402); on a stream, before the first byte. The list is materialized as a routing policy with an id derived from its content and scope, attributed in routing.policy_id and the x-merge-routing-policy-id header, and reusable later via routing_policy_id. model must be omitted or equal the first entry. Cannot be combined with routing_policy_id, a policy alias, or the top-level vendor / vendors pins. Also accepted on every SDK-compatible surface. See Sending a priority list with the request.
Tools the model may call: function tools, or provider-native tools passed through by type.
Tool choice policy: ‘auto’, ‘none’, ‘required’, or {‘type’: ‘function’, ‘function’: {‘name’: ’…’}}
Response format (json_object or json_schema structured output)
Optional project ID (UUID) to associate with this request
Include prompt-injection protection and DLP results for this request in a top-level guardrails object. In prompt-injection alert mode this also runs the direct scan on the request path (bounded by the alert-mode budget) so the score is available in the body; requests without the flag are unaffected.
Restrict routing to a specific vendor (hosting platform)
Processing tier: standard (default), flex, priority (fast is an alias), or ultrafast (OpenAI only), where default, auto, unspecified, and standard_only also mean standard; a tier the route doesn’t price returns 400 unsupported_params.
If the provider throttles the requested tier (429/503), retry once at the standard tier, billed at standard rates. The response service_tier shows the tier that actually served.
ID of a stored response (resp_...) whose conversation is replayed before this request’s input; an expired or unknown ID returns 404 response_not_found.
Set true to store this turn so a later request can continue it with previous_response_id; stored turns expire an hour after their last use and are never kept under zero data retention.
Number of most likely tokens to return per position; requires logprobs.
End-user identifier forwarded to the provider.
End-user identifier forwarded to providers that support it (OpenAI’s newer name for user).
Prompt-cache retention hint forwarded to OpenAI.
Whether the model may make several tool calls in one turn; omit to keep the provider default.
Output modalities to generate; include image for inline image generation on a route that supports image output, otherwise the request returns 400.
Top-level prompt-caching setting for providers that cache automatically.
Caching session hint, up to 256 characters, also accepted as the X-Session-Id header, which wins when both are set.
Extended thinking settings; the model can reason between tool calls.
Provider-specific options, keyed by provider family.
Multi-model fusion settings, used when model is fusion.
Response
The model that generated the response, in ‘provider/model-name’ format
Routing metadata (only present when include_routing_metadata=true)
Prompt-injection protection and DLP results (only present when include_guardrails_metadata=true). Also present on blocked 422 responses beside error.