AI Gateway Auto Router automatically selects an AI model for each request. It helps reduce costs while maintaining the quality of responses by matching requests with the model best suited to the task.
For each request, AI Gateway first builds a pool of eligible models. It excludes models that do not support the request format or its inputs. For example, if a request includes an image, only models that accept image input are eligible. It also accounts for credentials, billing configuration, and spend limits configured for the gateway. Unhealthy providers are excluded and restored automatically once they recover.
A classification step then analyzes the conversation to estimate the type and difficulty of the task. A scoring step weighs the expected quality of each remaining candidate against its cost, and ranks the models by fit. AI Gateway attempts the top-ranked model first, and falls back to the next one if a provider cannot serve the request.
To use Auto Router, send requests to your AI Gateway and specify cloudflare/auto in place of a provider-specific model. Refer to Unified API (OpenAI compat) for endpoint and authentication instructions.
curl -i -X POST "https://gateway.ai.cloudflare.com/v1/$CLOUDFLARE_ACCOUNT_ID/$CLOUDFLARE_GATEWAY_ID/compat/chat/completions" \
--header "cf-aig-authorization: Bearer $CLOUDFLARE_API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"model": "cloudflare/auto",
"messages": [
{
"role": "user",
"content": "hello"
}
]
}'The response headers show which model Auto Router selected:
cf-aig-routed-model: openai/gpt-5.6-luna
cf-aig-routing-reason: cost_optimal_within_pool
cf-aig-routing-decision-id: 91e0b970-33f0-4921-b8dc-127412e7b103To call Auto Router from a Worker, use the AI binding and specify the AI Gateway ID:
const response = await env.AI.run(
"cloudflare/auto",
{
messages: [
{
role: "system",
content:
"You are a helpful oracle powered by lava lamps. Answer clearly, with a flicker of dry humor.",
},
{
role: "user",
content: "hello",
},
],
},
{
gateway: {
id: "my-gateway",
},
},
);const response = await env.AI.run(
"cloudflare/auto",
{
messages: [
{
role: "system",
content:
"You are a helpful oracle powered by lava lamps. Answer clearly, with a flicker of dry humor.",
},
{
role: "user",
content: "hello",
},
],
},
{
gateway: {
id: "my-gateway",
},
},
);For multi-turn conversations, such as chats or coding agents, send a cf-aig-session-id header with each request. Auto Router then keeps the same model for an entire turn, so requests can take advantage of prompt caching. A turn is a user message plus any follow-up requests, such as tool calls. Supported clients, such as OpenCode, do this automatically.
Without a session ID, Auto Router selects a model for each request. Requests in the same conversation may go to different models, so they cannot reuse the prompt cache.
Auto Router detects turns from the message history. To mark turns yourself, send a cf-aig-turn-id header. To turn off session affinity, including for supported clients, send cf-aig-no-session-affinity: true.
At the start of each turn, Auto Router switches models only when the expected benefit outweighs the cost of losing the cache. For example, if a conversation moves from a simple question to a multi-file coding task, Auto Router may switch to a more capable model.
Auto Router supports the following request headers:
| Header | Description |
|---|---|
cf-aig-session-id |
Identifies the session that the request belongs to. Enables session affinity. |
cf-aig-turn-id |
Identifies the turn that the request belongs to. Overrides turn detection from message history. |
cf-aig-no-session-affinity |
Set to true to select a model for each request instead of pinning a model for the turn. |
cf-aig-allowed-models |
A comma-separated list of models that Auto Router can select from. Supports the * wildcard, for example anthropic/*. |
cf-aig-allowed-providers |
A comma-separated list of providers that Auto Router can select from, for example openai,anthropic. |
By default, Auto Router selects from a set of default models. cf-aig-allowed-models replaces the default models with the models you provide. Entries can match any model in Models, including additional models. The * wildcard does not match /, so use anthropic/* instead of *. If an entry matches no model, AI Gateway returns a 400 error. If you send both headers, a model must match both.
The response includes headers that describe the routing decision. These headers can help you identify the selected model, understand why it was selected, and correlate the request with AI Gateway logs for support and diagnostics:
| Header | Description |
|---|---|
cf-aig-routed-model |
The model that served the request, for example anthropic/claude-sonnet-5. |
cf-aig-routing-reason |
Why that model was selected. Refer to Routing reasons. |
cf-aig-routing-decision-id |
The routing decision identifier. |
cf-aig-request-id |
The request identifier. |
The cf-aig-routing-reason header has one of these values:
| Value | Description |
|---|---|
cost_optimal_within_pool |
Auto Router selected the model with the best balance of expected quality and cost. |
forced_by_candidate_pool |
Only one model was eligible, so no selection was needed. |
pinned_by_turn |
Auto Router reused the model selected earlier in the turn. |
fallback_candidate_unavailable |
The selected model could not serve the request, so AI Gateway used another eligible model. |
fallback_key_not_in_candidates |
The routing decision was not valid, so AI Gateway selected a model without it. |
fallback_router_error |
The router was unavailable, so AI Gateway selected a model without it. |
fallback_router_timeout |
The router did not respond in time, so AI Gateway selected a model without it. |
fallback_unsupported_input |
The request had no text to classify, so AI Gateway selected a model without the router. |
Auto Router selects from a pool of models managed by Cloudflare. The pool may change over time.
By default, Auto Router selects from the following models. To provide your own set of models, use the cf-aig-allowed-models or cf-aig-allowed-providers header.
| Model | Provider |
|---|---|
anthropic/claude-fable-5 |
anthropic |
anthropic/claude-opus-5 |
anthropic |
anthropic/claude-sonnet-5 |
anthropic |
openai/gpt-5.6-luna |
openai |
openai/gpt-5.6-sol |
openai |
openai/gpt-5.6-terra |
openai |
xai/grok-4.5 |
xai |
Auto Router only selects these models when they match cf-aig-allowed-models or cf-aig-allowed-providers.
Additional models
| Model | Provider |
|---|---|
@cf/deepseek-ai/deepseek-v4-flash-0731 |
workers-ai |
@cf/deepseek-ai/deepseek-v4-pro-0813 |
workers-ai |
@cf/google/gemma-4-26b-a4b-it |
workers-ai |
@cf/moonshotai/kimi-k2.7-code |
workers-ai |
@cf/openai/gpt-oss-20b |
workers-ai |
@cf/qwen/qwen3.8-27b |
workers-ai |
@cf/zai-org/glm-5.2 |
workers-ai |
alibaba/qwen3.8-max |
alibaba |
anthropic/claude-fable-5.1 |
anthropic |
anthropic/claude-opus-4.8 |
anthropic |
anthropic/claude-opus-5.5 |
anthropic |
anthropic/claude-sonnet-4.6 |
anthropic |
fireworks/glm-5.3 |
fireworks |
fireworks/glm-5.3-flash |
fireworks |
minimax/m3 |
minimax |
openai/gpt-5.5 |
openai |
openai/gpt-6-astra |
openai |
openai/gpt-6-luna |
openai |
openai/gpt-6-sol |
openai |
xai/grok-4.6 |
xai |
The Provider column uses the values that cf-aig-allowed-providers accepts.
OpenCode ↗︎ can use Auto Router as a custom model. Set the CLOUDFLARE_ACCOUNT_ID, CLOUDFLARE_GATEWAY_ID, and CLOUDFLARE_API_TOKEN environment variables, then add this opencode.json file to your project root:
{
"$schema": "https://opencode.ai/config.json",
"enabled_providers": ["cloudflare-auto-gateway"],
"provider": {
"cloudflare-auto-gateway": {
"npm": "@ai-sdk/openai-compatible",
"name": "Cloudflare AI Gateway",
"options": {
"baseURL": "https://gateway.ai.cloudflare.com/v1/{env:CLOUDFLARE_ACCOUNT_ID}/{env:CLOUDFLARE_GATEWAY_ID}/compat",
"apiKey": "",
"headers": {
"cf-aig-authorization": "Bearer {env:CLOUDFLARE_API_TOKEN}"
}
},
"models": {
"cloudflare/auto": {
"name": "Cloudflare Auto",
"modalities": {
"input": ["text", "image"],
"output": ["text"]
},
"limit": {
"context": 1000000,
"output": 128000
},
"options": {
"max_completion_tokens": 32000
}
}
}
}
}
}Start OpenCode with opencode and select Cloudflare Auto. For more OpenCode configuration options, refer to OpenCode.
The @cloudflare/aig-opencode-plugin ↗︎ plugin shows which model Auto Router selected for each request. It requires OpenCode 1.18.29 or later. To install it, run the command for your OpenCode version, then restart OpenCode:
opencode plugin add @cloudflare/aig-opencode-pluginopencode plugin @cloudflare/aig-opencode-plugin --globalThe plugin adds the following features:
- AI Gateway sidebar section. Shows the selected model, session ID, request ID, routing decision ID, and trace ID for the latest request in the current session. Select the section to copy these values.
/aig-showcommand. Prints the AI Gateway headers from the latest request, includingcf-aig-routing-reason, without sending anything to the model. Run/aig-show --helpfor options, such as showing earlier requests or including request and response bodies./aig-capturecommand. Run/aig-capture offor/aig-capture onto stop or resume capturing requests. Capture is on by default. Captured requests are kept in memory and cleared when OpenCode restarts.