Is your AI bot policy out of sync with robots.txt? Cloudflare's new Bot Preference Sync keeps them aligned automatically!
Hi, it's Shii! Today I've got a feature from the Cloudflare Blog that looks unassuming but could be a real win for anyone running a site!
Cloudflare BlogWhat was announced?
Cloudflare's Blog announced a new feature called Bot Preference Sync. It automatically reflects your AI crawler policy into your robots.txt file. Whatever you choose for the three categories, Search, Agent, and Training, on Cloudflare gets synced straight into your robots.txt.
The story so far
Until now, even if you set your AI bot policy on Cloudflare, you still had to maintain your robots.txt, a static file, separately on your own. As a result, it was easy to end up with mismatches, like robots.txt disallowing a specific crawler while your actual enforcement rules weren't blocking it. When your stated preferences and your enforced rules disagree, the crawler doesn't get the right signal either, so it was a bit of a waste.
What changes
Once you turn on Bot Preference Sync, Cloudflare's settings get automatically prepended to the top of your existing robots.txt. The Disallow directives you already had stay exactly as they were, so there's no risk of them getting overwritten and disappearing. Now you don't need to hand-edit robots.txt anymore, changing the policy in the dashboard is enough.
Dive Deep
There are three categories you can configure: Search, Agent, and Training.
- Search / Agent: choose from Allow, Block on pages that serve ads, or Block everywhere
- Training: a new Disallow option was added, and choosing it auto-generates a no-training entry in robots.txt
For example, if you choose Allow Search, Allow Agents, Disallow Training, a Disallow entry for training crawlers gets inserted at the top of robots.txt something like this.
# BEGIN Cloudflare Bot Preference Sync
User-agent: TrainingBot1
User-agent: TrainingBot2
User-agent: MixedUseBot-Extended
Disallow: /
# END Cloudflare Bot Preference Sync
For crawlers that do both search and training, Cloudflare also lays out transparency requirements they need to meet:
- Respect the no-training setting in robots.txt
- Offer an option to hide AI summaries
- Provide URL-level usage and metrics
- Publicly confirm there's no impact on traditional search results
It's available to every customer, from the Free tier all the way up to Enterprise. New customers get it on by default, and existing customers will see a migration prompt over the coming weeks. If your site is monetized with ads, you can choose Disallow Training as your default option.
Wrap-up
- Bot Preference Sync automatically syncs Cloudflare's AI bot policy into robots.txt
- Previously, dashboard settings and robots.txt could drift out of sync
- Search / Agent offer Allow, Block on ad pages, or Block everywhere; Training now adds a Disallow option
- Your existing Disallow entries stay intact, with Cloudflare's settings prepended at the top
- Available to every customer from Free to Enterprise; new customers get it on by default, existing customers roll out over the coming weeks
If you've been hand-maintaining robots.txt or want your AI crawler policy to actually stay consistent, this update is for you!