Skip to main content

ENTRY 016 · AI-MODE · By Answer Engineered Research

-18.8ppvs+8.8pp

18.8 Points Lost, 8.8 Recovered. Only One Was Measured Cleanly.

AI MODE FORCED -18.8ppOVERVIEWS HIDDEN +8.8pp

A preregistered experiment on live Google Search put causal numbers on AI Mode and AI Overviews. Mid-study, a Google HTML change broke one arm of the test.

· 6 MIN

What the experiment actually did

Wang, Gleason, Bart, Wilson and Metaxa recruited US-based participants in waves between 17 and 19 March 2026, screening for people who used Chrome as their primary browser and Google as their primary search engine. Recruitment came from Prolific and from Northeastern students in a work-study program with one of the authors.

A total of 1,444 people enrolled and installed the extension. Of those, 1,100 made at least one search during the seven-day treatment window, and 956 completed the post-experiment survey — 343 in No AI Search, 304 in Current Search, 309 in AI Mode Search.

Three conditions:

  • Current Search — unmodified Google. AI Overviews may appear, AI Mode is reachable.
  • No AI Search — the extension hides AI Overviews.
  • AI Mode Search — every search is redirected into AI Mode.

That design is the point. Everyone measuring AI search from the outside is comparing sites, or comparing time periods, and hoping nothing else moved. This compares randomly assigned people using their own Google, in their own browser, in the same week.

The two numbers, with their intervals

Condition, vs unmodified GoogleEffect on click-through95% CI
AI Mode Search (forced)-18.8 pp-22.2 pp to -15.3 pp
No AI Search (Overviews hidden)+8.8 pp+2.3 pp to +15.3 pp

Both intervals exclude zero, so both effects are real in the statistical sense. But look at the widths. The AI Mode interval spans about 7 points. The No AI interval spans 13 — it is nearly twice as uncertain, and its lower bound sits at 2.3 points, close enough to nothing to matter.

The paper’s own summary of the AI Mode arm: “relative to standard Google Search, enforcing AI Mode reduced external click-through by 18.8 percentage points, reduced daily search sessions, and reduced the fraction of users clicking through to news sites, Reddit, and Wikipedia alike.”

Why the second number is softer than it looks

This is in the paper, in the results, in the authors’ own words. It has to survive into any honest use of the finding.

The No AI Search condition depended on the extension recognising an AI Overview and hiding it. During the study period, Google shipped a change to AI Overview HTML that broke the extension’s logic for identifying them. The result, quoted directly:

“Overall, 51.1% of AI Overviews were successfully hidden and the median participant in this condition had 50% of AI Overviews hidden.”

So in the arm meant to show you a Google without AI Overviews, the median participant still saw half of them.

The authors did the correct thing rather than the convenient thing. Their preregistration already specified that they would report local average treatment effect estimates if compliance failed, and they had defined individual compliance as receiving the assigned treatment on at least 90% of searches. When that definition collapsed, they switched to measuring compliance continuously and estimated the effect with two-stage least squares, instrumenting exposure by assignment, with HC1 robust standard errors.

That is a defensible repair. It also means the 8.8-point figure is a different kind of estimate than the 18.8-point figure: it is scaled up from a partially delivered treatment, not read off a treatment that landed. Putting them in the same sentence with the same confidence is the mistake waiting to be made.

There is a second-order point here that nobody will make. A vendor’s measurement tooling can be broken by a silent markup change on Google’s side, mid-measurement, without anyone being told. That happened to a team of academics who were watching for it and who published the damage. It is happening continuously to commercial AI-visibility trackers that have no comparable incentive to disclose it.

AI Mode did not simply move clicks. It made people like search less.

The traffic result is the one publishers will read. The perception results are the ones that complicate Google’s position.

On a standardised scale, being pushed into AI Mode reduced satisfaction by 0.73 standard deviations (95% CI: -0.88 to -0.59). It increased time spent: minutes per session rose by 0.43 minutes (95% CI: 0.20 to 0.66). And it pushed people toward the exits — the paper reports an 11.2% increase in competitor search use, with AI Mode users expressing substantial intent to switch to Bing if AI Mode became mandatory.

The abstract’s own framing: AI Mode “reduces click-through rates and erodes user experience and trust in information found on Google.”

Note what that sentence does not license. Removing AI Overviews did not improve trust. The paper: “We do not find evidence that exposure to No AI Search impacted trust [95% CI: -0.10, 0.38].” A null result with an interval straddling zero. If you want to argue that stripping AI Overviews makes users trust Google more, this experiment does not give you that, and the gap between the two arms is again wider than the headline pairing suggests.

What this cannot tell you

The authors list their own limits, and they are the limits that matter for anyone applying this to a site.

The sample is not America. The analysis sample skewed 75% under 45, 87% with at least some college, and 58% Democrat against 20% Republican. Effects “may differ in magnitude or direction in a more representative sample.”

Seven days is short. The paper: the design “allows us to detect short-term behavioral and attitudinal shifts but not longer-run adaptation — it is possible that some effects would attenuate, or that new ones would emerge, over a longer exposure period.” People forced into an unfamiliar interface for a week dislike it. That is not the same finding as people disliking it in a year.

The treatment is forced, not chosen. The AI Mode arm redirected every search into AI Mode. Nobody’s actual Google works that way. This is the causal effect of AI Mode being mandatory, which is a real question about where Google is going, but it is not a measurement of AI Mode as it exists today.

And the data is five months old. Recruitment ran 17 to 19 March 2026. The preprint went up on 18 August. Every number here describes a version of Google Search that has since been iterated on repeatedly — including, as the paper itself documents, at least one markup change shipped mid-study.

The part worth keeping

Strip out the framing and one thing survives that no correlational study has produced: a randomised, preregistered, causal estimate that AI Mode moves publisher click-through by roughly eighteen points in the direction publishers feared, on live Google, with real users.

That is a genuinely important result and it did not need the second number attached to it.

The second number will get attached anyway, because 18.8 against 8.8 reads like an asymmetry with a moral in it. The honest version is narrower: one arm of this experiment measured what it set out to measure, the other arm measured half of it and was scaled up with a method the authors preregistered in advance for exactly this failure. Both belong in the record. Only one of them belongs in a headline.

Sources

  • Wang, Gleason, Bart, Wilson, Metaxa, “AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence,” arXiv:2608.18352, submitted 18 August 2026. Abstract · Full text

Every figure above is quoted from the preprint’s own results and limitations sections. We did not open the replication data.