Skip to content

AI Gateway

Last updated View as MarkdownAgent setup

Every AI Search instance is connected to a Cloudflare AI Gateway. External-provider embedding and reranking calls run through your connected gateway. Query rewriting and response generation also run through your gateway.

Workers AI embedding and reranking usage is included in AI Search usage. These calls do not appear in your AI Gateway logs or analytics and are not billed separately as Workers AI usage.

To choose or change which gateway your instance uses, see Models.

Observe your model calls

AI Gateway records the model requests routed through your gateway. This includes all query-rewriting and response-generation calls, plus external-provider embedding and reranking calls.

  • Analytics: Track the number of requests, tokens used, cost, latency, and errors across your model calls.
  • Logs: Inspect individual requests and responses, including the effective system prompt, rewritten queries, and generated answers.

Use models from other providers

By default, AI Search uses Workers AI models. To use models from other providers, such as OpenAI or Anthropic, add your provider keys to AI Gateway and select those models in AI Search.

  1. Add your provider keys with Bring Your Own Keys.
  2. Connect the gateway and select the models in your AI Search settings. For details, see Models.

Guard against unsafe content

Use AI Gateway Guardrails to screen the prompts and responses that flow through your instance and block content that is unsafe or inappropriate. To detect and handle sensitive information, such as personal or financial data, use Data Loss Prevention (DLP).

Improve resilience

Configure request retries and model fallbacks so that a model call can automatically retry or fall back to another model when a provider returns an error.

Caching and rate limiting

Some AI Gateway features act on every request that passes through your gateway. These features can interfere with external-provider embedding and reranking, query rewriting, and response generation.

Do not turn on AI Gateway caching for the gateway connected to your AI Search instance. This matters most for embedding requests. AI Search relies on fresh embeddings to build its vector index and to match each query against it, so serving cached embeddings can store or return incorrect vectors and quietly degrade the accuracy of your search results. To cache search results, use AI Search's own Similarity cache instead.

Similarly, avoid setting rate limiting on this gateway. Rate limits apply to AI Search's own model calls, including the many embedding requests made while indexing, and can interrupt indexing and querying.

Was this helpful?