ENTRY 017 · AI-CRAWLERS · By Answer Engineered Research
3vs0
Cloudflare Set Four Conditions for Googlebot. It Names Nobody Who Meets Them.
Bot Preference Sync lets a mixed-use crawler keep search access on a training-disallowed site by meeting four disclosure conditions. Each is self-attested.
What Cloudflare has done, and what it has said it will do
These two get merged in almost every write-up of this story, so they are worth separating.
Cloudflare has done this: on 1 July 2026 it split AI bot management into three categories a site owner can set independently, called Search, Agent and Training, on every plan including Free. On 21 August it published the four disclosure conditions and announced Bot Preference Sync, a feature that writes your dashboard policy into your robots.txt so that, in its phrasing, “the preference you set is the preference you publish.”
Cloudflare has said it will do this: set new defaults on 15 September 2026, in these terms. “For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.” On the same date, multi-purpose crawlers “will be allowed/blocked according to all of their behaviors”. And Bot Preference Sync itself “will be available to all customers, on every plan, in the coming week” — written on 21 August, with no exact date given.
None of the September items has happened. This post does not predict what will happen on 15 September, and you should be suspicious of anything that does. What exists today is a published policy, a published set of conditions, and a dashboard section.
One detail is easy to miss, and it is the reason this reaches sites that made a decision a long time ago. The 15 September change to multi-purpose crawlers applies to customers who selected to block training “either through the new options to manage AI traffic, or through the legacy Block AI bots service”. The old one-click block is in scope.
The four conditions, in Cloudflare’s own words
Quoted verbatim from the 21 August post, in the order they appear there:
- “The bot must respect, via any mechanism, a ‘no training’ preference in robots.txt”
- “They give site owners a way to opt out of AI summaries.”
- “They provide URL-level visibility into which pages were made available for training, as well metrics on search results, so you can see how your content was used for search and for training.”
- “They can show publicly that Disallowing Training does not hurt your traditional search results.”
Counted against each other, the two posts produce this:
| What Cloudflare’s two posts name | Count |
|---|---|
| Disclosure conditions a mixed-use crawler operator must meet | 4 |
| Crawlers named as multi-purpose and in scope on 15 September | 3 |
| Operators named as currently meeting the four conditions | 0 |
| Tests, auditors or evidence standards described for those conditions | 0 |
Named as multi-purpose: 3. Named as meeting the conditions: 0.
The last row needs a qualifier, because it is a count of an absence. Cloudflare does say where compliance is published: “Bots of leading AI models and service providers that meet these criteria are tracked publicly in the AI bot transparency section in Cloudflare Radar, which includes examples in which best practices are honored, as well as when they are not.” That is a venue for results. It is not a description of how a result is reached, and the post contains no such description.
Self-attested is the whole mechanism
Cloudflare is explicit about who supplies the evidence: “the owners of bots that perform both Search and Training will need to provide additional information in order to not be blocked when ‘Disallow Training’ is set.”
Provide. The operator provides.
Work through what each condition would take to check independently, and the reason that verb matters becomes clear.
Condition 1 concerns what an operator does with data after it has collected it. An edge network can see a request. It cannot see whether those bytes ended up in a training run, because nothing in the request distinguishes a crawl that indexes from a crawl that trains. That is the entire premise of the mixed-use problem Cloudflare set out to solve, restated as a condition.
Condition 2 is the most externally observable of the four. Either a documented way to opt out of AI summaries exists or it does not. Even here the post does not say who judges whether a particular control counts as one.
Condition 3 asks for URL-level reporting on which pages were made available for training. Whether such a report exists is checkable from outside. Whether it is complete is not.
Condition 4 has no stated standard at all. The phrase “can show publicly that Disallowing Training does not hurt your traditional search results” leaves open what counts as showing, what counts as hurt, and what the comparison is against.
Cloudflare’s own framing of the deal is direct: “this is a way of making Transparency the price of admission.” The price is a disclosure. The disclosure is the operator’s account of itself.
The public list we could not read
The compliance list is the one thing that would settle this story, so we went to get it.
Cloudflare Radar’s AI Insights page, at radar.cloudflare.com/ai-insights, is where the post points. A plain fetch of it on 25 August 2026 was refused outright. A fetch carrying an ordinary browser user agent returned the page. Inside it, the section is real: a heading reading “AI bot transparency” and a description reading “Tracking best practices from bots of leading AI models and service providers”. Directly under that description is a link back to the Bot Preference Sync post.
Below the heading, the served HTML contains a figure marked as a widget in a loading state, with a row of placeholder bars where a table belongs and no operator names anywhere in the document. The table renders client-side. Fetching the page does not get you its contents.
So we cannot tell you who is on that list. Not that it is empty. We do not know, and neither does anyone reporting on this who has not opened it in a browser.
That gap is the whole story in one place. If Google, Apple or Microsoft appear on that list, the four conditions have been exercised by the operators they were written for, and a training-disallow setting changes very little about search indexing. If they do not appear, the escape hatch exists on paper and has not been walked through. Those are two different worlds, the list is what separates them, and the list is the part we could not read.
The setting that actually matters before 15 September
If you are a Cloudflare customer who has ever switched on a training block, including through the legacy Block AI bots service, the 1 July post gives one concrete instruction: “if a website owner wants to opt out of these new default configurations, they can easily mark this in their Security settings any time leading up to September 15, which will confirm that they want no changes on Training crawlers that also crawl for Search purposes.”
That is the action. Marking it means multi-purpose crawlers are not swept into your training block. Leaving it unmarked means your training block is enforced against them under what Cloudflare calls “the most restrictive applicable rules”.
Three more things in the same posts are worth knowing before you touch the dashboard.
Bot Preference Sync writes to your robots.txt, and Cloudflare says existing content survives: “If a site owner already has a robots.txt file, the contents added by Bot Preference Sync will be prepended to the existing material, so any existing Disallow directives are maintained.” Read the file after it lands anyway.
It is on by default for new customers. Existing customers on the legacy managed robots.txt feature are prompted to review and confirm rather than switched silently. That prompt is the last look you get at the policy before it is published on your domain.
It ignores your custom rules. Cloudflare says so plainly: “Because Bot Preference Sync is designed to tackle policy decisions made category-wide rather than case-by-case, it will not directly read from individual custom rules with more complex logic.” If you have a private arrangement with one operator expressed as a custom rule, the generated robots.txt will not reflect it, and what you publish will disagree with what you enforce. That is the exact failure the feature was built to remove.
What this evidence cannot tell you
One vendor is describing its own product. Every fact above comes from two Cloudflare blog posts and one Cloudflare dashboard. Cloudflare writes the conditions, classifies the crawlers, operates the block, and publishes the compliance list. There is no second party in this story, and no independent reading of any of it.
The scale figure is self-reported. Cloudflare’s own line is “the more than 20% of web domains that sit behind Cloudflare”, published in Cloudflare’s own post with no methodology attached. Domains is not defined there, and registered domains, active sites and zones are three different counts.
The classification is one company’s reading of three others. Calling Googlebot, Applebot and BingBot multi-purpose crawlers is Cloudflare’s account of how Google, Apple and Microsoft run their crawlers. We looked and did not find a statement from any of the three confirming or disputing it.
Nothing here has been measured. There is no traffic data in this story, no before-and-after, no experiment, no sample. It is a policy document and a feature announcement, which is a weaker class of evidence than the causal work on AI search we covered on 21 August, and a different class again from the crawl and referral logs in Cloudflare’s own AEO dashboard.
And the deadline has not arrived. Every consequence discussed here is conditional on a default Cloudflare has announced and not yet applied.
The conditions are published, the deadline is dated, and the list that reconciles them was still loading.
Sources
- Cloudflare, “Say it once: introducing Bot Preference Sync”, 21 August 2026 — https://blog.cloudflare.com/bot-preference-sync/
- Cloudflare, “Your site, your rules: new AI traffic options for all customers”, 1 July 2026 — https://blog.cloudflare.com/content-independence-day-ai-options/
- Cloudflare Radar, AI Insights, section “AI bot transparency” at radar.cloudflare.com/ai-insights — fetched 25 August 2026. Not linked above because a plain fetch of it is refused and only a browser-agent request returns the page; no figure from it is used here.
Both blog posts were fetched and read in full on 25 August 2026 and returned HTTP 200. Every quotation above is verbatim from those two pages, including the phrasing of condition 3.