Fusion
Fusion sends one prompt to a panel of models at once, and a judge model merges their answers into one response. Use it when answer quality matters more than latency or cost: several mid-tier models reconciled by a judge often beat any one of them.
How fusion works
Each panel model answers the whole prompt on its own; fusion doesn’t split it into sub-tasks. A panel model that errors, returns empty text, or is blocked by policy is dropped, and the request succeeds if at least one answers. The judge (synthesis_model) reads the conversation and the surviving answers and writes the final one.
Make a fusion request
Set model to fusion and add a fusion block to a POST /v1/responses request. Fusion works only on the native Responses API, without streaming.
max_tokens applies to each call separately, so leave the judge room to write the merged answer.
Every underlying call is billed and logged as its own request, so a three-model panel costs about four calls.
Add web search
Add {"type": "merge:web_search"} to tools and each panel model runs its own web search loop, so panelists can find different sources; the judge doesn’t search. Each panelist’s citations are in its fusion.candidates[].annotations. The final answer keeps a citation only when the judge’s text contains its URL, so paraphrased sources are dropped. Web search is the only tool fusion accepts, since the panel and judge run inside Gateway and can’t hand a function call back to you.
Reference
The fusion block
Response
output is the judge’s answer in the standard /v1/responses shape, and usage sums every panel call and the judge. The fusion object describes the run: