You can upload files directly to an instance or connect an external data source.
| Data Source | Description |
|---|---|
| Built-in storage | Upload files directly to an instance. Available by default on every instance. |
| Website | Connect a domain you own to index website pages. |
| R2 Bucket | Connect a Cloudflare R2 bucket to index stored documents. |
For website data sources, Parse types covers how AI Search finds the pages to index.
AI Search can ingest a variety of file types. The following plain text files and rich format files are supported.
| Format | File extensions | MIME type |
|---|---|---|
| Text | .txt, .rst |
text/plain, text/x-rst |
| Log | .log, .log.gz |
text/plain |
| Config | .ini, .conf, .env, .properties, .gitignore, .editorconfig, .toml |
text/plain, text/toml |
| Markdown | .markdown, .md, .mdx, .mdoc |
text/markdown |
| LaTeX | .tex, .latex |
application/x-tex, application/x-latex |
| Script | .sh, .bat, .ps1 |
application/x-sh, application/x-msdos-batch, text/x-powershell |
| SGML | .sgml |
text/sgml |
| JSON | .json |
application/json |
| XML | .xml |
application/xml |
| SQL | .sql |
application/sql |
| YAML | .yaml, .yml |
application/x-yaml |
| CSS | .css |
text/css |
| JavaScript | .js |
application/javascript |
| TypeScript | .ts, .tsx |
application/typescript |
| PHP | .php |
application/x-httpd-php |
| Python | .py |
text/x-python |
| Ruby | .rb |
text/x-ruby |
| Java | .java |
text/x-java-source |
| C | .c |
text/x-c |
| C++ | .cpp, .cxx |
text/x-c++ |
| C Header | .h, .hpp |
text/x-c-header |
| Go | .go |
text/x-go |
| Rust | .rs |
text/rust |
| Swift | .swift |
text/swift |
| Dart | .dart |
text/dart |
| Terraform | .tf |
text/plain |
| EMACS Lisp | .el |
application/x-elisp, text/x-elisp, text/x-emacs-lisp |
AI Search uses Markdown Conversion to convert rich format files to markdown. The following table lists the supported formats that will be converted to Markdown:
Format | File extensions | Mime Types | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
PDF Documents |
|
| ||||||||||||
Images 1 |
|
| ||||||||||||
HTML Documents |
|
| ||||||||||||
XML Documents |
|
| ||||||||||||
Microsoft Office Documents |
|
| ||||||||||||
Open Document Format |
|
| ||||||||||||
CSV |
|
| ||||||||||||
Apple Documents |
|
|
1 Image conversion uses two Workers AI models for object detection and summarization. See Workers AI pricing for more details.
AI Search applies these file size limits:
| File type | Maximum size |
|---|---|
| PDF with OCR enabled | 10 MiB |
| PDF without OCR | 4 MiB |
| Plain text, code, configuration, markup, and other formats listed in Plain text file types | 10 MiB |
| Other formats converted to Markdown | 4 MiB |
Files that exceed these limits are not indexed. They appear in the error logs.
AI Search supports .jpg, .jpeg, .png, .webp, .gif, .bmp, .tif, .tiff, .heic, and .heif images. Multimodal embedding models embed supported images directly. Text-only embedding models create captions before embedding images.
Optical character recognition (OCR) extracts text from scanned PDFs and images. Set indexing_options.use_ocr to true when creating or updating an instance. OCR is disabled by default and changing this setting triggers a full reindex.
OCR is available for every account. OCR usage is billed as image-processing ingestion tokens. For pricing details, refer to Limits and pricing.