Multimodal
Chat models on Gateway accept images, documents, audio, and video as input, and dedicated endpoints generate images, speech, transcripts, and video. Find your modality below, and check the model catalog for models that support it.
Support matrix
Each compat surface takes its own wire shape (OpenAI image_url, input_audio, and file parts, Anthropic source blocks, AI SDK file parts), translated to the served vendor’s format. Support is per route, so check your vendor’s vendors.<vendor>.capabilities on GET /v1/models. A route that can’t take a modality you sent returns 400 unsupported_params with param: input before any vendor call.
Input blocks
Native content blocks on /v1/responses:
Each cap is the decoded inline total for that modality across the request. URL media doesn’t count, so send long video as a URL. Over a cap, chat requests fail with 413 payload_too_large; the dedicated media endpoints have their own limits.
Audio input is reported as usage.audio_input_tokens and billed at the route’s audio rate when it has one; other media bills as ordinary input tokens.