Set rejectIfBusy when your application should not wait in a capacity queue. Workers AI rejects the synchronous inference request if capacity is unavailable.
For the native REST API, add rejectIfBusy to the request options object:
curl --request POST \
--url "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/ai/run/@cf/google/gemma-4-26b-a4b-it" \
--header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"messages": [
{
"role": "user",
"content": "Explain what a capacity queue is."
}
],
"options": {
"rejectIfBusy": true
}
}'For the Workers AI binding, pass rejectIfBusy in the third argument to env.AI.run():
const response = await env.AI.run(
"@cf/google/gemma-4-26b-a4b-it",
{
messages: [
{
role: "user",
content: "Explain what a capacity queue is.",
},
],
},
{ rejectIfBusy: true },
);const response = await env.AI.run(
"@cf/google/gemma-4-26b-a4b-it",
{
messages: [
{
role: "user",
content: "Explain what a capacity queue is.",
},
],
},
{ rejectIfBusy: true },
);Do not add rejectIfBusy to the model input object. The binding only applies this option from the third argument.
For OpenAI-compatible Chat Completions, add options at the top level of the request body:
curl --request POST \
--url "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/ai/v1/chat/completions" \
--header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"model": "@cf/google/gemma-4-26b-a4b-it",
"messages": [
{
"role": "user",
"content": "Explain what a capacity queue is."
}
],
"options": {
"rejectIfBusy": true
}
}'OpenAI clients that preserve custom fields can send this option. Clients that remove unknown fields do not apply it, so requests proceed normally.
Rejected requests return HTTP status 429 and internal error code 3040. The error message is Capacity temporarily exceeded, please try again.
Refer to Workers AI errors for error details.