Self-hosted models

Route to model endpoints you run yourself, alongside Merge's curated vendors

Connect inference endpoints you run yourself, such as a vLLM cluster or a dedicated Baseten deployment, to Gateway. Your apps call them through the same Gateway API as every other model, and you can add them to routing policies next to curated vendors. Use them for fine-tuned or private models, for capacity you’ve already paid for, or when prompts must stay on your own infrastructure.

How it works

You register an endpoint once: its URL, its API key, and the models it serves. When a request names one of those models, Gateway sends it to your endpoint and returns the response in the same format as any other model.

  • Only your organization can see, call, or route to your endpoints.
  • Your API key is write-only. Gateway encrypts it and never shows it again in the dashboard or API.
  • Merge doesn’t bill this traffic. You pay for your own infrastructure; nothing is charged to your Gateway balance.
  • Zero data retention (ZDR) organizations can use them, because prompts go to infrastructure you control. Failover to curated vendors still only uses ZDR-compliant vendors.

Endpoint requirements

Your endpoint must:

  • Serve an OpenAI-compatible chat completions API (/v1/chat/completions, which vLLM, TGI, SGLang, Baseten, and most inference servers support)
  • Use HTTPS and be reachable from the public internet. Private and loopback addresses are rejected.
  • Accept its API key as a bearer token (Authorization: Bearer <key>)

Add an endpoint

You need the Manage credentials permission to add, edit, or delete an endpoint.

  1. In the Gateway dashboard, open Configure and go to Self-hosted models
  2. Choose Add endpoint
  3. Enter a Display name, and a Slug (lowercase letters, numbers, and hyphens). The slug can’t be changed later.
  4. Enter the Endpoint URL, for example https://llm.example.com/v1, and the API key
  5. Add each model the endpoint serves:
    • Model ID: the name your applications will use to call it
    • Upstream name: only if your endpoint uses a different name than the Model ID
    • Context window, Max output tokens, and whether it supports Streaming, Tool calling, and Vision
    • Input price and Output price (optional, $ per 1M tokens). Merge never bills this price.
  6. Choose whether to Allow failover to Merge curated vendors (see Failover)
  7. Choose Add endpoint

The endpoint is saved as a Draft. Drafts can’t receive traffic until they pass a test and you activate them.

Test and activate an endpoint

Test sends a few tiny requests to your endpoint and shows a result for each check. It doesn’t change anything.

CheckWhat it catches
Endpoint reachableA wrong URL, a DNS or TLS problem, or a timeout
API key acceptedA wrong or expired key (the endpoint answers 401 or 403)
OpenAI-compatible responseA reply that isn’t in the chat completions shape
Token usage reportedA reply with no token counts
Model servedA wrong Upstream name, checked for each model
Streaming worksA server that ignores streaming or sends a broken stream, checked for each model that declares streaming

Activate runs the same checks and turns the endpoint on only if all of them pass. Otherwise it stays a draft and the results show what to fix. If someone edits the endpoint while the checks run, activation stops, so what goes live is always what passed.

Deactivate stops traffic to an active endpoint right away without deleting it. You can activate it again later.

Edits to an active endpoint apply immediately, without a new test, so you can rotate a key with no downtime: edit the endpoint and enter the new key (leave the field blank to keep the current one). Run Test after changing the URL or key.

Call a self-hosted model

Call the model by the Model ID you declared, exactly as you entered it:

curl https://api-gateway.merge.dev/v1/responses \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "llama-3.3-70b-instruct",
"input": [
{ "type": "message", "role": "user", "content": "Summarize this ticket in one sentence." }
]
}'

Your self-hosted models also appear under Self-hosted in the model picker, and in GET /v1/models for your organization.

Model IDs can be a plain name (llama-3.3-70b-instruct) or include a prefix (acme/llama-3.3-70b-instruct). If you declare an ID that’s also in Merge’s catalog, your endpoint takes precedence: calls from your organization that name it always go to your endpoint, whatever the curated vendors charge.

Use self-hosted models in routing policies

Add self-hosted models to a Priority policy from the Self-hosted group in the model picker, next to curated models. Gateway tries models in the order you list them and moves on to the next entry if your endpoint fails.

Failover

Allow failover to Merge curated vendors (on by default) controls what happens when your endpoint fails:

  • On: in a Priority policy, Gateway tries the next model in the list, including another self-hosted endpoint. And if your Model ID is also in Merge’s catalog, a request your endpoint can’t serve (server error, timeout, or unreachable) can go to a curated vendor for that model, billed at its normal rate.
  • Off: Gateway returns your endpoint’s error. Prompts never leave your infrastructure.

A rejected API key is never sent to a curated vendor for the same model. You get the error back, so a bad key shows up right away.

Limits

  • Test doesn’t check tool calling or vision. Gateway trusts what you declare.
  • Slugs can’t be changed. To rename an endpoint, delete it and add it again.
  • Apps and policies that use an old Model ID stop reaching the model if you change it
  • Up to 50 models per endpoint
  • If two endpoints declare the same Model ID, Gateway always uses the same one, and you can’t pick which

Next steps