Text to speech
Synthesize audio from text
POST /v1/audio/speech turns text into audio for narration, voice replies, or notifications. The response body is the raw audio, with no JSON wrapper; Content-Type names its format.
Voices
The Gemini TTS models run on the google vendor only and take up to 8,192 input tokens.
Reference
Other fields are forwarded to the vendor. Speech bills per input character (unit: per_character on GET /v1/models). X-Merge-Billed-Characters gives the count billed after any DLP redaction, and X-Merge-Vendor names the vendor.