Text to speech

Synthesize audio from text

POST /v1/audio/speech turns text into audio for narration, voice replies, or notifications. The response body is the raw audio, with no JSON wrapper; Content-Type names its format.

curl https://api-gateway.merge.dev/v1/audio/speech \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/tts-1", "input": "Your order has shipped and arrives Thursday.", "voice": "alloy"}' \
-o speech.mp3

Voices

ModelVoicesOutput
openai/tts-1, openai/tts-1-hdOpenAI voices such as alloy and novaresponse_format, MP3 by default
google/gemini-3.8-flash-tts, google/gemini-3.8-flash-lite-ttsGemini voices such as Kore or Puck, or omit voice. OpenAI names fail with No matching speaker voice found.WAV, whatever response_format says
bland/bland-speech-v3OpenAI voice names, mapped to Bland voices, or a Bland voice UUIDWAV
Other vendors, such as minimax/speech-2.6-turboThe vendor’s own voice IDsVendor-specific

The Gemini TTS models run on the google vendor only and take up to 8,192 input tokens.

Reference

FieldNotes
modelRequired. A canonical model ID, since @alias/... returns 400 alias_not_supported.
inputRequired. The text to speak.
voiceModel-specific, see above
response_formatmp3, opus, aac, flac, wav, or pcm
speed0.25 to 4.0
instructionsVoice direction, on models that support it
vendorPin the execution vendor
customerCustomer UUID to scope the key, budget, and usage to

Other fields are forwarded to the vendor. Speech bills per input character (unit: per_character on GET /v1/models). X-Merge-Billed-Characters gives the count billed after any DLP redaction, and X-Merge-Vendor names the vendor.

Next steps