Cache-aware routing
Cache-aware routing
When a conversation’s turns land on different vendor routes of a model, each switch rewrites the prompt cache instead of reading it. Cache-aware routing keeps a session’s follow-up requests on the vendor that served it while that cache is warm. It’s on for every organization, with no configuration.
How a session is identified
First match wins:
- Explicit session. The
X-Session-Idheader (wins) or bodysession_id. On the/v1/openaichat completions and responses endpoints,prompt_cache_keyalso counts, the same hint automatic caching uses. - Thread or harness session. An
X-Merge-Thread-Idheader, or else a coding agent’s own session header:x-claude-code-session-id,session-idorsession_id(Codex CLI), orx-session-affinity. - Conversation opener. A hash of the first system and first user messages.
Affinity is tracked per requested model or policy, so a session that mixes models tracks each one and switching policies starts fresh. It applies on /v1/responses and every compatible endpoint.
When affinity applies
Affinity only moves the remembered vendor to the front of the routing order. It never pins, never fails a request, and adds negligible latency. Ordinary vendor selection takes over when:
- The remembered vendor is unavailable or filtered out for the request
- Its cache window (per route, 10 minutes by default) has passed
- The session now resolves to a different model
- Reading that vendor’s cache is no cheaper than its own input price or the cheapest input price on the model’s other routes
- The route’s cache pricing isn’t known
It has no effect on a model with one vendor route, or when caching never engaged. On explicit caching-family routes, such as Anthropic and Claude on Bedrock, caching needs a cache_control marker on the stable prefix (the system prompt or earlier turns). Prompt caching lists each vendor’s family.