Fusion

Send one prompt to several models at once and merge their answers into a single response

Fusion sends one prompt to a panel of models at once, and a judge model merges their answers into one response. Use it when answer quality matters more than latency or cost: several mid-tier models reconciled by a judge often beat any one of them.

How fusion works

Each panel model answers the whole prompt on its own; fusion doesn’t split it into sub-tasks. A panel model that errors, returns empty text, or is blocked by policy is dropped, and the request succeeds if at least one answers. The judge (synthesis_model) reads the conversation and the surviving answers and writes the final one.

Make a fusion request

Set model to fusion and add a fusion block to a POST /v1/responses request. Fusion works only on the native Responses API, without streaming.

curl https://api-gateway.merge.dev/v1/responses \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "fusion",
"input": [
{"type": "message", "role": "user", "content": "Explain the CAP theorem in plain terms."}
],
"fusion": {
"analysis_models": ["openai/gpt-5.4-mini", "anthropic/claude-sonnet-4-6", "deepseek/deepseek-v4-pro"],
"synthesis_model": "anthropic/claude-sonnet-4-6"
}
}'
CallReceives
Each panel modelmax_tokens, temperature, top_p, stop, tools, tool_choice, tags, project_id, include_routing_metadata
JudgeOnly max_tokens, tags, project_id, include_routing_metadata

max_tokens applies to each call separately, so leave the judge room to write the merged answer.

Every underlying call is billed and logged as its own request, so a three-model panel costs about four calls.

Add {"type": "merge:web_search"} to tools and each panel model runs its own web search loop, so panelists can find different sources; the judge doesn’t search. Each panelist’s citations are in its fusion.candidates[].annotations. The final answer keeps a citation only when the judge’s text contains its URL, so paraphrased sources are dropped. Web search is the only tool fusion accepts, since the panel and judge run inside Gateway and can’t hand a function call back to you.

Reference

The fusion block

FieldRequiredDescription
analysis_modelsYesAt least two distinct model IDs (vendor/model or a bare name). Duplicates are removed
synthesis_modelNoThe judge. Defaults to the first analysis_models entry
synthesis_instructionsNoReplaces the judge’s default instruction, for example to force a format
include_candidatesNoDefaults to true. false returns an empty candidates list

Response

output is the judge’s answer in the standard /v1/responses shape, and usage sums every panel call and the judge. The fusion object describes the run:

FieldDescription
synthesis_modelThe judge that wrote the answer
models_requestedThe panel after duplicates were removed
models_succeededHow many panel models contributed
candidates[]Each panel model’s model, vendor, status (ok or error), text, usage, error when it failed, and annotations when web search ran

Errors

StatusCodeCause
400fusion_config_requiredmodel is fusion but there’s no fusion block
400fusion_stream_unsupportedstream is true
400fusion_unsupported_toolstools contains something other than web search
400fusion_routing_policy_unsupportedrouting_policy_id is set. The panel is chosen by analysis_models
400empty_inputinput is empty
400web_search_unavailable, server_tool_zdr_unavailable, unsupported_web_search_engineThe web search tool can’t run as configured
422Validation erroranalysis_models has fewer than two distinct models
503Vendor errorEvery panel model failed, or the judge failed

Next steps