Skip to main content

ENTRY 015 · AI-CRAWLERS · By Answer Engineered Research

150:1227:1

179:1, 6,000:1, 50,000:1: Three Tools, One Metric, No Shared Definition

150:1227:1Q1 2026Q2 2026

Scrape-to-referral is the metric of the moment. TollBit, Microsoft Clarity and Cloudflare each publish a version of it, and none of them count the same thing.

· 6 MIN

Where each number comes from

SourceNumberWhat it coversPublished
TollBit179:1European publishers on TollBit’s networkState of the Bots, 2026 Q1 & Q2
Microsoft Clarity6,000:1A figure in the product screenshot, not the post text13 August 2026
Cloudflare118:1 to ~50,000:1Bots from leading AI companies, observed on Cloudflare’s networkAttribution Business Insights

TollBit’s is the most specific claim of the three, and it is the only one stated as a sentence in the report rather than shown in an interface:

“It takes 179 AI bot visits to get a single human visitor referral in return from AI applications for European publishers — a rate over 3x worse than North American sites.”

The same report says the exchange got worse across the first half of the year: “the scrape-to-referral ratio went from 150:1 in Q1 2026 to 227:1 in Q2 2026.” So even inside one dataset, one methodology and one six-month window, the headline figure moves by half again depending on which quarter you quote.

Cloudflare’s range is wider than TollBit’s entire dataset:

“Bots from leading AI companies have been observed with a range of crawl-to-referral ratios: we noted ratios of 118:1 up to nearly 50,000:1 around the time of our Content Independence Day in 2025.”

What each one actually counts

This is the part that decides whether the numbers can be compared, and it is usually left out of the coverage.

Cloudflare counts referrals through UTM parameters. Its own developer documentation defines the metric as “the number of crawls sent by this company, vs. the number of visitors who visit you through a referral link from that company, tracked through UTM parameters.” A referral that arrives without a UTM tag is not in the denominator.

Clarity counts through its own analytics tag, joined to its Bot Analytics data, and it restricts the calculation by domain coverage. Microsoft’s own caveat in the announcement: “Ratios are calculated only across correctly mapped domains, helping prevent misleading rollups when CDN and Clarity coverage do not fully align.”

TollBit counts inside its own publisher network. The report attributes its figures to “Data from publishers on TollBit.”

Three different denominators, three different populations, three different collection layers. A site could plausibly sit in all three products at once and be handed three ratios that differ by two orders of magnitude, without any of the three being wrong.

The most-quoted number in the story is a screenshot

We read Microsoft’s announcement in full. It does not contain a ratio.

The post describes what the card does — “Compare AI scrape activity against referral traffic in a single metric, helping you understand whether AI platforms are returning value in the form of visits rather than only extracting content” — and lists the features. There is no benchmark figure, no sample size, no date range, and no cross-site average anywhere in the text.

The 6,000:1 that travelled through the coverage comes from the illustrative product image. Search Engine Land’s Barry Schwartz described it plainly as a reading of the screenshot: “This seems to say that this AI tool sent 41 referrals to your website, and if you do the math, it has a 6,000 scrapes-to-1-referral ratio based on its crawl activity related to referral activity” (Search Engine Land, 13 August 2026). ppc.land made the same point: “That figure is a product screenshot rather than a published benchmark, and Microsoft does not present it as a cross-site average.”

Both outlets labelled it correctly. It still ended up in circulation as though Microsoft had measured something.

Who is in the sample

TollBit’s numbers come from publishers who pay TollBit. That is not a random sample of the web; it is a set of publishers who were already worried enough about AI scraping to buy tooling for it. The direction of that bias is genuinely unclear — such sites may be more targeted, or simply better instrumented — but it is not a general population, and the report does not claim to be one.

TollBit also sells the fix for the problem it measures, and the report’s foreword is credited to a publisher executive rather than a neutral party. None of that makes the numbers wrong. It does mean the number and the remedy come from the same place, which is worth saying out loud, because the coverage we read did not say it.

Cloudflare is in the same position from the other direction: it sells bot management, and its ratio range is drawn from traffic crossing its own network.

The one figure in TollBit’s report that limits it hardest is TollBit’s own: “these numbers reflect only identified bots, so the true scale of AI scraping is likely even higher.” Unidentified scrapers are not in the numerator.

The denominator is undercounted, and that cuts the other way

Every one of these ratios has the same structural weakness, and it pushes all of them in the same direction: referrals are systematically undercounted, so the ratios read worse than the underlying exchange.

ppc.land’s analysis sets out the mechanism: Google confirmed in May 2025 that noreferrer elements in AI Mode link code were stripping referrer values and causing those clicks to register as direct traffic, and OpenAI added UTM parameters to ChatGPT links in June 2025. Neither fix was complete. As ppc.land puts it: “A ratio built on an undercounted denominator overstates the imbalance.”

That is ppc.land’s reasoning, not Microsoft’s or Cloudflare’s, and we are attributing it rather than asserting it. But it follows directly from Cloudflare’s own published definition. If your denominator is UTM-tagged visits, then every AI referral that arrives untagged makes your ratio look more extractive than it is.

What this can and cannot tell you

It cannot tell you the real scrape-to-referral ratio for your site. Nothing here establishes a true value. Three vendors measured three populations three ways and got three answers.

It cannot tell you which tool is most accurate. We have no ground truth to check them against, and neither does anyone else. That is the actual problem.

It does not mean the underlying imbalance is invented. All three products point the same way: AI crawlers take a great deal more than they send back. The direction is consistent across independent methodologies, which is a real signal even when the magnitudes are not comparable. TollBit’s within-dataset quarter-over-quarter move — 150:1 to 227:1 on one consistent methodology — is more informative than any cross-vendor comparison, precisely because the method held still.

What it does tell you is that “scrape-to-referral ratio” is not yet a measurement. It is a product feature name that three companies adopted independently, each defining it against whatever data they already had.

The quick version

If someone quotes you a scrape-to-referral ratio, the number is close to meaningless until you know three things: which population it was measured over, how referrals were detected, and whether it came from a report sentence or an interface screenshot.

Ask those three questions before you repeat the figure. On the most widely shared number in this story, the answer to the third one is: a screenshot.