ENTRY 019 · AI-CRAWLERS · By Answer Engineered Research
3vs1
Half of Web Traffic Is Non-Human. Three Cloudflare Numbers Became One.
Cloudflare says more than 50% of internet traffic is non-human. An explainer read that as AI bots. Three Cloudflare percentages, three different denominators.
The two sentences, side by side, both fetched
We fetched both documents ourselves on 31 August 2026. Both returned HTTP 200.
The Conversation piece is by Dana McKay, Associate Dean, Interaction, Technology and Information at RMIT University, and Damiano Spina, Senior Lecturer in RMIT’s School of Computing Technologies. Its own metadata carries the date: pubdate reads 20260830, og:updated_time reads 2026-08-30T20:16:47Z. The claim quoted above sits mid-article, introduced by the line “This change in traffic patterns isn’t a small or hypothetical problem.”
The Cloudflare sentence it summarises sits under the heading “The agentic Internet is here”, in a report titled “Content Independence Day, one year on: building the business model for the agentic Internet”, published 1 July 2026 at 13:00 UTC and last modified 15 July 2026.
The gap between them is one word doing an enormous amount of work. “Non-human” is a classification of what sent the request. “AI bots” is a classification of what the request was for. The first contains the second, and it contains other things too.
Cloudflare’s own post separates the categories the sentence merges
You do not have to reach outside the Cloudflare post to see that non-human is the wider box. The post does the separating itself, in the section headed “Crawlers have changed their purpose”:
“When looking at the crawlers Cloudflare identifies by purpose, the composition of crawler traffic tells the story clearly: 52% of crawler requests are now for AI training as of June 2026, up from 22% in Spring 2025. Mixed-use crawlers (those blending search, agent use, and training) represent over 36% of activity. Pure search crawling now represents a small and declining share of overall crawler activity, despite remaining critical for publisher visibility.”
If 52% of crawler requests are for AI training, then 48% of crawler requests are for something else, and Cloudflare names one of those something-elses in the next breath: pure search crawling, which it says remains “critical for publisher visibility”. Crawlers are themselves a subset of non-human traffic, and the report treats agents and crawlers as distinct populations throughout.
So the 52% is the AI-specific number in this report, and it is a percentage of crawler requests, not of all traffic on the internet. It is also the number that has actually moved: 22% in Spring 2025 to 52% in June 2026. That is a real, large, dated shift, and it is the finding a reader would want. It is not the finding the sentence reports.
The three denominators, laid out
| Figure | What it is a percentage of | Cloudflare’s own sentence | Date |
|---|---|---|---|
| more than 50% | all traffic on the Internet, classified as non-human | ”This year, agent traffic crossed a historic threshold for the first time: more than 50% of traffic on the Internet is now non-human.” | published 1 July 2026, modified 15 July 2026 |
| 52% | crawler requests, classified by purpose as AI training | ”52% of crawler requests are now for AI training as of June 2026, up from 22% in Spring 2025.” | measurement dated June 2026, baseline Spring 2025 |
| over 36% | crawler activity, classified as mixed-use | ”Mixed-use crawlers (those blending search, agent use, and training) represent over 36% of activity.” | published 1 July 2026, modified 15 July 2026 |
All three sentences are verbatim from the same page, fetched on 31 August 2026. Reading down the middle column is the whole finding: 3 denominators in the source document, 1 in the sentence that summarises it.
The number 36% appears twice in that post, meaning two unrelated things
This is worth naming, because it shows how ordinary the compression is. The “over 36% of activity” above is a share of crawler traffic. Later in the same report, in a section headed “Part III: A unique view of the ecosystem”, Cloudflare writes: “More than 20% of the web sits behind Cloudflare’s network. Of the world’s most-visited websites, 36% rely on our network, and more than 40% of the Fortune 500 are Cloudflare customers.”
That second 36% is a share of websites, not of traffic. Same page, same number, no relationship. A document that packs five differently-based percentages into a few screens is easy to compress wrongly, and doing so requires only that the denominators be dropped — which is what happens to almost every statistic that travels.
The page the claim links does not contain the words “AI bot”
The hyperlink under “estimates over half of all web traffic” points at Cloudflare Radar’s worldwide traffic dashboard for the last 12 months. We fetched it. HTTP 200.
It is a JavaScript dashboard, so a plain fetch returns the page shell rather than the live chart values, and we make no claim about what number it currently displays. The shell does contain the page’s own labels, and those are checkable. The relevant sections are headed “Bot traffic worldwide” and “Bot vs. Human worldwide”. The description under the second reads “Bot (automated) vs. human HTTP requests distribution to HTML content”. The note in that chart card’s footer adds a further scope: “Percentage of HTTP requests classified as bot (automated) or human. Filtered to HTML responses, representing web page traffic.”
That footer note also carries a fourth scoping detail nobody repeats downstream: whatever the chart shows, it is a share of requests for web pages, not of all traffic of every kind.
The string “AI bot” does not appear anywhere in the HTML that page served us. Not in a heading, not in a chart description, not in a footer note.
Cloudflare does publish an AI-specific bot dashboard. It is a different Radar page, and its first section is headed “AI bot & crawler traffic”. The general and the AI-specific views are two separate surfaces on the same product, and the sentence links the general one while making the AI-specific claim.
The “30% or more of the top 10,000 sites” figure carries no citation
The same sentence opens with a second statistic: that Cloudflare “manages 30% or more of the top 10,000 sites on the internet”. In an article with 23 outbound links in its body, this claim carries none. We looked for its source and did not find one, so we are not repeating it as a fact and you should not either until somebody produces the document.
What Cloudflare’s own report says about its footprint, in the passage quoted above, is “More than 20% of the web sits behind Cloudflare’s network” and “Of the world’s most-visited websites, 36% rely on our network”. Neither is a percentage of the top 10,000 sites, and neither is 30%. We are not claiming the article’s figure was derived from either sentence. We are recording that we could not source it at all.
The September 15 inference does not follow from Cloudflare’s own wording
Further down, in a section headed “What’s happening in the short term”, the article says: “Come September 15, Cloudflare sites will block AI crawlers by default on pages that contain advertising (and therefore make money for content creators). This means up to 30% of the world’s top sites will no longer appear in Google AI Overviews summaries.”
The first sentence links to a TechCrunch write-up. The second links to nothing.
We did not use the TechCrunch article. We went to Cloudflare’s own announcement of that change, published 1 July 2026, and read the policy in Cloudflare’s words:
“On September 15, 2026, we’ll be setting new defaults for each of these three classifications. For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.”
And, on the same page:
“Of course, customer choice is paramount: if a website owner wants to opt out of these new default configurations, they can easily mark this in their Security settings any time leading up to September 15, which will confirm that they want no changes on Training crawlers that also crawl for Search purposes.”
Four things are in Cloudflare’s own wording that are not in the summary of it. The new default applies to new domains onboarding to Cloudflare. It applies on pages that display ads. It blocks two of the three categories, Training and Agent, while Search “will remain allowed by default”. And a site owner can opt out of it in a settings page before the date.
A further mechanism on the same page affects crawlers that do more than one job at once: multi-purpose crawlers “will be allowed/blocked according to all of their behaviors”, and Cloudflare names three, saying that “multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training”. That rule is real and consequential, and it turns on a choice the customer already made. We took it apart in an earlier post on the four disclosure conditions attached to it.
What none of this supports is a number for how many of the world’s top sites will drop out of AI Overviews. That would require knowing how many top sites are new Cloudflare onboardings, how many serve ads on the relevant pages, how many have already chosen to block Training, how many opt out before the date, and how Google’s sourcing responds to each. The article states the conclusion without flagging any of those steps as inference. We are not making the opposite prediction either: any percentage published for 15 September before 15 September is arithmetic on assumptions.
One statistic in the same article checks out exactly
Later in the piece: “one recent study found that already, around 1 in 6 sources used by AI search tools is itself an AI-generated website.” That sentence links to an arXiv paper, and we opened it.
The paper is “Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources”, submitted 22 May 2026 by Mowafak Allaham and Nicholas Diakopoulos. Its abstract describes “an audit of four generative search engines (ChatGPT, Copilot, Gemini, Perplexity) using a total of 712 real-world human-generated queries”, and reports “evidence of AI-generated sources being cited across all four generative search engines (~16% of cited sources)”.
Around 1 in 6 is 16.7%. The paper says approximately 16%. That is a fair paraphrase of a linked primary source with the study’s scope intact, in the same article, by the same two authors, published the same day. Whatever went wrong earlier in the piece is not a property of the piece as a whole, which is why this is a denominator problem and not a credibility problem.
What this evidence cannot tell you
We did not read the live values on the Radar dashboard. It renders its charts in the browser, and our fetch returned the static shell. Every claim we make about that page is about its labels and its served HTML, not about the number a visitor sees on the chart.
We cannot tell you how the compression happened, and we did not try. A sentence that merges three denominators is a methodology failure with many ordinary causes, and we have no evidence about which one applies. Nothing above should be read as a claim about anybody’s care or intent.
We did not source the “30% or more of the top 10,000 sites” figure, in either direction. We could not find a Cloudflare statement of it. That is not the same as establishing it is wrong.
We are not forecasting 15 September, and we did not fetch or rely on the TechCrunch article the article links there. The policy is published, the date is set, and the outcome is unmeasured. Everything we say about it comes from Cloudflare’s own page, quoted verbatim.
Cloudflare’s figures are Cloudflare’s own network measurements, and we have no independent check on any of them. More than 50% non-human, 52% of crawler requests, over 36% mixed-use: each is one vendor’s classification of traffic crossing its own network, published in a report arguing for a market in which that vendor sells the metering. That does not make the numbers wrong. It does mean all three denominators are defined by one interested party, and a second measurement of any of them would be worth more than a louder restatement of the first.
Sources
- Dana McKay and Damiano Spina, “AI is eating website traffic, websites are blocking AI – and reliable information is getting harder to find”, The Conversation, 30 August 2026 — https://theconversation.com/ai-is-eating-website-traffic-websites-are-blocking-ai-and-reliable-information-is-getting-harder-to-find-289817
- Cloudflare, “Content Independence Day, one year on: building the business model for the agentic Internet”, published 1 July 2026, modified 15 July 2026 — https://blog.cloudflare.com/agentic-internet-bot-report/
- Cloudflare, “Your site, your rules: new AI traffic options for all customers”, 1 July 2026 — https://blog.cloudflare.com/content-independence-day-ai-options/
- Cloudflare Radar, worldwide traffic, last 12 months — the page the claim links — https://radar.cloudflare.com/traffic?dateRange=52w
- Cloudflare Radar, AI Insights — the AI-specific dashboard on the same product — https://radar.cloudflare.com/ai-insights
- Mowafak Allaham and Nicholas Diakopoulos, “Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources”, arXiv, submitted 22 May 2026 — https://arxiv.org/abs/2605.23684
All six URLs were fetched on 31 August 2026 and every one returned HTTP 200. Every quotation above is verbatim from the source named beside it.