Image generation

Generate images with a dedicated endpoint or inline in chat

Dedicated image models such as openai/gpt-image-2 and xai/grok-imagine-image generate through POST /v1/images/generations. Chat models that also output images, such as google/gemini-3.1-flash-image, generate inline on /v1/responses. To change an existing image, use image editing.

Dedicated endpoint

curl https://api-gateway.merge.dev/v1/images/generations \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-image-2", "prompt": "a lighthouse at dawn, watercolor", "size": "1024x1024"}'

Images come back as base64 in data[].b64_json, with images_generated, cost in USD, and usage on token-priced models. The X-Merge-Vendor header names the vendor. For xAI, Gateway requests base64 when you omit response_format, since xAI refuses URL output for zero data retention accounts; an explicit "url" is forwarded.

Inline in chat

Set modalities: ["text", "image"] with a model that lists image in capabilities.output (other routes return 400 unsupported_params with param: modalities). Images arrive in output[].content[] as "type": "image" blocks with base64 data.

cURL
curl https://api-gateway.merge.dev/v1/responses \
-H "Authorization: Bearer $MERGE_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemini-3.1-flash-image",
"input": [{"type": "message", "role": "user", "content": [{"type": "text", "text": "draw a small blue star"}]}],
"modalities": ["text", "image"]
}'

Reference

FieldNotes
modelRequired. A canonical model ID, since @alias/... returns 400 alias_not_supported.
promptRequired
n1 to 10
size, quality, style, output_formatModel-specific
response_formatb64_json or url
vendorPin the execution vendor
customerCustomer UUID to scope the key, budget, and usage to
provider_optionsVendor-specific options. Other top-level fields are also forwarded.

pricing.unit on GET /v1/models says how a model bills: per_token meters image output tokens (usage.image_output_tokens), so cost scales with resolution, and per_image charges a flat rate per image, with optional tiers in output_per_image_by_resolution.

Next steps