AI Search Adds Custom Domains, Access Control, and Cross-Instance Search, Plus Sitemap-less Crawling!
Hi, I'm Shii! Today's news comes straight from the Cloudflare Changelog. Cloudflare's "AI Search" service just picked up a bundle of features for putting your search endpoint in front of real users. Let's take a look!
Cloudflare ChangelogWhat was announced?
On the Cloudflare Changelog, Cloudflare announced a new update to AI Search. AI Search is a service that gets you from a data source to a working search endpoint quickly. This release adds what you need to put that endpoint in front of real users.
Specifically, four new options were added.
- Serving your public endpoint from a custom domain
- Restricting access with Cloudflare Access
- A namespace-level endpoint that searches across several instances at once
- A new crawl mode for sites that don't have a complete sitemap
Importantly, the previous behavior stays the default. The announcement is explicit that every one of these is opt-in, so nothing changes until you turn it on.
The story so far
A "public endpoint" in AI Search is a URL that a site or app can query directly, with no authentication in front of it. Until now, that URL was always a generated hostname under search.ai.cloudflare.com. You couldn't serve it from your own domain, you couldn't restrict who could query it, and there was no way to search across multiple instances (each instance being one data source) from a single URL.
Indexing had a gap too. Sites were only crawled through their sitemap, so if a site's sitemap was out of date or incomplete, pages you wanted indexed could be missing.
What changes
This release brings together what you need to take AI Search from a working prototype to something you can actually put in front of users.
- Serve from a custom domain: You can now serve your public endpoint from a hostname in a zone you own, like
https://search.example.com/search - Restrict with Cloudflare Access: Once your endpoint is on your own domain, you can put Cloudflare Access in front of it, so for example you can limit the
/mcppath to specific agents instead of anyone who finds the URL. Agents authenticate with an Access service token, and people opening it in a browser sign in through your identity provider - Search across instances via namespace: A namespace can expose its own public endpoint with
/search,/chat/completions, and/mcppaths that fan out across the instances you choose - Crawl sites without a sitemap: Website data sources support a new
discoverparse type that starts at your source URL and collects pages from both your sitemaps and the links it finds while crawling
If you've wanted to turn a site into a search engine and actually ship it to users, this closes a lot of the gap between "it works" and "it's live."
Dive Deep
Here's a closer look at the examples from the changelog.
Querying multiple instances at once:
curl https://ns-{NAMESPACE_ENDPOINT_ID}.search.ai.cloudflare.com/search \
--header "Content-Type: application/json" \
--data '{
"messages": [{ "content": "How do I configure AI Search?", "role": "user" }],
"ai_search_options": { "instance_ids": ["docs", "support"] }
}'
You list the instance IDs you want to search in ai_search_options.instance_ids, so a single request can cover multiple data sources, like docs and support together.
Configuring the discover parse type:
curl -X POST "https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai-search/instances" \
-H "Authorization: Bearer {API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"id": "my-ai-search",
"type": "web-crawler",
"source": "example.com",
"source_params": {
"web_crawler": {
"parse_type": "discover",
"discover_options": { "source": "links", "limit": 5000, "depth": 3 }
}
}
}'
discover_options takes source (where to discover pages from), limit (a cap on how many pages to collect), and depth (how far to follow links from the source URL); the changelog's example sets limit to 5000 and depth to 3. Since pages get picked up by following links during the crawl, not just from a sitemap, this helps if your sitemap has fallen behind what's actually on your site.
On the Access side, the interesting part is that you can apply different authentication to different paths on the same custom domain: service tokens for the /mcp path your agents call, and identity-provider sign-in for people opening the endpoint in a browser. That combination fits the common need of wanting a public-facing search endpoint without leaving it open to anyone who finds the URL.
For the full configuration steps for custom domains, Access, and namespaces, see the AI Search documentation.
Wrap-up
- AI Search's public endpoint can now be served from a custom domain instead of only a generated hostname
- You can pair it with Cloudflare Access to restrict paths like
/mcpto specific agents or users - A namespace-level public endpoint lets you search across multiple instances from a single URL
- The new
discoverparse type crawls pages that aren't listed in your sitemap - All of these are opt-in additions, previous behavior stays the default
If you've been wanting to turn your site or docs into a real search engine you can put in front of users and teammates, this update is exactly what you were waiting for!