Web search

Give models access to current web information through Gateway

Web search lets a model look up current information on the public web mid-request and answer with URL citations. Use it when the answer needs recent facts or sources that aren’t in your prompt. It doesn’t search your own data or documents.

Add {"type": "merge:web_search"} to tools. The model decides whether to search, Gateway runs the queries and returns results to it, and the model answers. Gateway tells the model to treat results as untrusted, ignore instructions in them, and cite URLs.

curl https://api-gateway.merge.dev/v1/responses \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-sonnet-4-6",
"input": [
{"type": "message", "role": "user", "content": "What changed in the latest Kubernetes release? Include sources."}
],
"tools": [
{"type": "merge:web_search", "parameters": {"max_results": 5, "allowed_domains": ["kubernetes.io"]}}
]
}'

The tool also works on the OpenAI SDK chat completions and Responses endpoints, the AI SDK surface, and with stream: true (the first token arrives after any search finishes). It isn’t available in batch. To require a search, set tool_choice to {"type": "merge:web_search"}. After the first search, the model may answer.

Read citations and usage

Text output carries url_citation annotations, one per source URL. Parsing links from the text instead misses sources the model didn’t restate.

{
"type": "text",
"text": "The latest release adds...",
"annotations": [
{"type": "url_citation", "url": "https://kubernetes.io/blog/release-notes", "title": "Kubernetes release notes", "content": "Relevant excerpt"}
]
}

Native /v1/responses, the OpenAI SDK Responses endpoint, and the AI SDK surface use this flat shape. /v1/openai/chat/completions returns message.annotations in OpenAI’s nested url_citation shape and streams them on deltas. The response’s usage.server_tool_use reports web_search_requests and web_search_results.

Choose an engine

Leave engine at auto unless you need a specific index. auto uses the first available engine in this order: Parallel, Exa, You.com, Interfaze, Tavily, Browserbase. With zero data retention on, it skips Browserbase, and pinning Browserbase returns 400 server_tool_zdr_unavailable.

A timeout, reset, 429, or 5xx is retried once on the same engine, then the search fails over to the other engines in the order above, or to your fallback_engines list (always on Merge-managed keys). Each search gets at most three attempts. Pinning engine or passing api_key turns off automatic failover, but an explicit fallback_engines list still applies. If every attempt fails, the model gets a tool error and answers from what it has.

The model can search several times per request. After six model turns, Gateway runs one last turn with tools off and attaches a server_tool_iteration_limit_reached warning.

OpenAI’s and Anthropic’s own search tools have a separate setting at Configure → Request settings → Defaults → Web search. Override it per request with provider_options.web_search set to auto, merge, or native.

SettingBehavior
Automatic (recommended)Runs the vendor’s search when the route can (OpenAI models on OpenAI, Anthropic models on Anthropic or Vertex AI, never Bedrock), otherwise runs it as merge:web_search with a web_search_remapped warning
Search through MergeAlways uses merge:web_search
Built-in search onlyReturns 400 web_search_requires_provider when the route can’t run it

Reference

Parameters

Set these in the tool’s parameters object.

ParameterDefaultDescription
engineautoauto, parallel, exa, youcom (alias you), interfaze (aliases openwebsearch, ows), tavily, or browserbase
max_results5Results per search, 1 to 25. Interfaze caps at 10 and Tavily at 20
max_total_resultsmax_results × 5Cap across all searches in the request, 1 to 250. Past it the model gets a tool error instead of another search
search_context_sizeEngine’s adaptive snippetslow (5,000 characters per result), medium (15,000), or high (30,000)
allowed_domains, excluded_domainsNoneHostnames to include or exclude. Browserbase filters after fetching, so it can return fewer results
user_locationNoneOnly country is used. Tavily and Browserbase ignore it
timeout_seconds10Timeout per search call
api_keyMerge-managed keyYour own key for the selected engine
fallback_enginesAutomaticEngines to try in order when the primary fails

Engines

EngineZero data retentionPrice per searchResults includedEach extra resultNotes
Parallel (parallel)Yes$0.00510$0.001Accepts an objective from the model. low, medium, high map to turbo, basic, and advanced (default advanced)
Exa (exa)Yes$0.00710$0.001
You.com (youcom)Yes$0.005AllNone
Interfaze (interfaze)YesVaries, $0.007 fallbackAll (up to 10)NoneBilled at the cost the service reports per call
Tavily (tavily)Yes0.008,or0.008, or 0.016 at highAll (up to 20)Nonehigh uses advanced search
Browserbase (browserbase)No$0.007None$0.001 per fetched pageReturns full pages as Markdown instead of snippets (default cap medium). Unfetched pages aren’t charged

Web search is billed separately from tokens, with your organization’s usual margin applied. A search that fails over is priced at each engine’s own rate.

Errors

StatusCodeCause
400unsupported_web_search_engineUnknown engine
400web_search_unavailableNo key available for the engine, or a fallback_engines entry has no Merge-managed key
400server_tool_zdr_unavailableA pinned or fallback engine doesn’t support zero data retention
400duplicate_server_toolMore than one web search tool in the request
400invalid_server_tool_parameters, invalid_search_context_sizeA parameter is invalid
400web_search_requires_providerBuilt-in search only is set and the route can’t run the vendor’s search

Next steps