Skip to main content
  • Home
  • Tech
  • Website Traffic in the Age of AI Crawlers: What Is Lost When the Answer Replaces the Click

Website Traffic in the Age of AI Crawlers: What Is Lost When the Answer Replaces the Click

Picture

Member for

1 year 3 months
Real name
SIAI Editor
Bio
SIAI Editor

Modified

AI answers replace clicks while crawlers take publishers' content
Blocking removes rigorous sources and weakens answer reliability
Verifiable crawlers and auditable payment are the minimum fix

In March 2025, the Pew Research Center recorded the behavior of 900 US adults in 68,879 Google searches and confirmed something that many publishers were already seeing in their statistics. When an AI summary appeared at the top of the page, users clicked on some result in 8 percent of visits, versus 15 percent when there was no summary and the link within the summary itself was opened by just 1 percent. The decline in website traffic is usually described as a change of medium, with readers moving toheir questions from the search engine to apps like ChatGPT and Gemini and with Google in the role of the main loser. The data gathered over the past two years paints a broader picture. The loss does not stop at the search engine, it passes to the content creators and finally reaches the credibility of the information that the audience reads.

How Website Traffic is Lost Between the Answer and the Crawler

Two mechanisms work in parallel. The first concerns demand: the response produced by a language model meets the reader's need before they need to visit any page. The IAB Tech Lab, the body that coordinates technical standards for digital advertising, estimates that AI summaries reduce publisher traffic by 20 percent to 60 percent, with losses of up to 90 percent on specialized websites and estimates lost advertising revenue at about $2 billion. The second mechanism concerns supply, i.e. the crawlers that collect the content on which these answers are based. In traditional search, the two sides were tied, since Google crawled a page to display a link that someone would click. Cloudflare counted how many pages each platform crawls for each visit it sends back and for one week of June 2025 the ratio for Anthropic was about 70,900 to 1. The company noted that the number may be overstated because mobile apps don't always state the source of the referral, but even with a generous correction the distance from the old search logic remains huge.

Figure 1: AI answer engines can use publisher content without returning the reader.

Content collection also comes at a cost that is not reflected in estimates of advertising revenue. Since the fall of 2025, websites of all sizes have started recording thousands of zero-time visits from Lanzhou, an industrial city in northwest China, with some of the traffic passing through Singapore. According to the platform Analytics.usa.gov, the city accounted for 14.7 percent of visits to US federal government websites in one period. Who is behind this move remains open, because the geographical attribution of an IP address shows where the infrastructure is registered, not who uses it. For the small publisher, the consequences are tangible. The statistics by which it decides what to write and how to sell it are distorted and advertising networks that see hundreds of thousands of empty sessions have removed websites from their programs, considering that they inflate the numbers. The phenomenon is not only Chinese. Similar waves are emerging from data centers in Virginia, where under a suburban name, declared crawlers of American companies, undeclared agents and monitoring tools are mixed.

Answers Read without Their Source

If replacing the click always gave correct answers, the issue would be mainly financial. The largest evaluation of AI assistants in information was coordinated by the European Broadcasting Union led by the BBC and published in October 2025. Twenty-two public broadcasters from 18 countries, in 14 languages, reviewed more than 3,000 responses of ChatGPT, Copilot, Gemini and Perplexity to current affairs questions. 45 percent had at least one major problem, 31 percent had serious problems in source attribution and 20 percent had serious accuracy errors, while for Gemini the percentage of responses with a significant problem reached 76 percent. Columbia University's Tow Center had come up with something similar with a different method when it asked eight search tools to locate the title, publisher and date of articles from snippets and found errors in more than 60 percent of queries. Paid versions of some tools were wrong with more certainty than free ones.

OpenAI researchers provided an explanation in September 2025 that helps show why errors persist. In their analysis, hallucinations arise from the statistical pressures of training and are maintained because most evaluations score in a way that rewards speculation more than the assumption of uncertainty. The scarcity of data on a topic increases errors without being their root. A language model produces the most likely continuation of a text and this property is improved by retrieving sources and better calibration without disappearing. Google's results page also had many weaknesses, but in front of a list of links, the reader chose which to open, saw the name of the publisher and decided, even roughly, whether a medical question was better answered by a hospital or by a forum. In front of an answer, the choice has already been made by the system and the percentage of 1 percent who return to the source shows how rarely the check is done.

Who Gains and Why Publishers Block Crawlers

A simple tally shows more than one loser and model providers are not necessarily on the winners' side. According to shareholder documents cited by The Information, OpenAI consumed $3.7 billion in cash flow in the first quarter of 2026, with revenue of $5.7 billion over the same period. Each answer has a computation cost that in much of the free use is not covered by revenue and the margin accumulates upstream, in hardware. Nvidia reported revenue of $130.5 billion for the fiscal year ended January 2025, of which $115.2 billion came from data centers. In this chain, those who sell chips and infrastructure are in a clearly better position than the editor of a blog or the company that subsidizes the answers by betting on future profitability.

Publishers mostly lack control. A website can't decide whether its text will be used in a response that falsifies it, has no way to ask for a correction and often doesn't even know if and when its content was detected. The robots.txt file is a request that no one is obligated to respect and in Google's case, the Google-Extended option excludes a text from Gemini training without excluding it from AI Overviews summaries. To stay out of them, a publisher would have to disappear from search. Under these circumstances, blocking became the dominant strategy of major news organizations. Analysis published by the Press Gazette in January 2026 found that 79 percent of the top news sites in the UK and US blocked at least one training crawler and 71 percent blocked trackers that pulled content in real-time.

Here the loss of traffic and credibility are tied in a loop. The sources that invest the most in checking their data are the ones that withdraw first, because they have the resources and the bargaining power to do so, so the answers are increasingly based on what is left open, on reposts, on copies and on pages made for search engines. As for the traffic recorded in Lanzhou or Singapore, there is no counterparty to negotiate with there and no authority to complain to. The publisher resorts to blocking entire networks, at the cost that along with the robots it cuts off any real reader who happens to pass through the same providers.

Open Weights, Payments and the Limits of Today's Solutions

The strongest counter-argument deserves to be said clearly. Summaries save time, give access to explanations that would otherwise require hours of reading and click-free search existed long before language models, since Pew itself recorded that about two-thirds of searches ended up unvisited in a result regardless of the summary. The argument correctly describes some of the value to the user. It leaves unanswered that in queries where the user would click somewhere, the summary almost doubles the loss and that the value the user gets is generated with material that someone else paid to exist.

Of the solutions being discussed, open-weight models return public access, since researchers and public bodies can review and host them, but they do not pay any authors or reduce errors. Licensing agreements pay for those with large files. The settlement in the Bartz v. Anthropic case, amounting to $1.5 billion, corresponds to about $3,000 for each of about 500,000 books, an amount that arose because a court requested it and that an independent editor has no way to claim. The most advanced step toward payment at the page level was taken by Google. According to Digiday, in September 2026 the company scaled up a pilot program within Search Console that pays when a piece of content contributes significantly to Gemini, AI Overviews and AI Mode responses. Participants see a monthly amount with no explanation of the calculation and participation is by invitation only.

The changes needed are distributed among different bodies. Companies that develop models can, without legislation, publicly separate their crawlers by purpose, make their identities verifiable, give each site access to its record of crawls and open a remediation channel for claims attributed to a specific source. Google and those experimenting with payments could publish the valuation model and open the programs to any site that meets minimum criteria. Small publishers have a reason to organize themselves into collective licensing bodies, as the music industry once did. Public intervention has a place where no company has an incentive to move first, such as in a registry of crawlers with mandatory identification and binding exemption statements in a machine-readable format.

Figure 2: A workable AI–publisher exchange depends on identity, permission and contribution.

When 1% Becomes the Rule

The 1 percent of visits that return from a summary to its source summarize a lot of the above. The move from the results page to the answer page kept the cost of the calculation, which is currently paid by loss-making companies and their investors and implicitly removed the mechanism that covered the cost of content through advertising and subscriptions. Creators lose readers and the ability to count those who remain, readers get answers with double-digit error rates and the more thoughtful sources are withdrawn from the material that feeds the models. Google's program is the first pay-per-contributor program, on terms it sets itself. Without verifiable crawler identity, binding exclusion statements and a valuation that can be checked by third parties, any such scheme will function as a concession rather than an obligation. Whether it will evolve into something more will be seen by when and if, the formula with which the payment is calculated will be published.


This article reflects the analytical judgment of the author and does not constitute policy advice or the official position of any affiliated institution.

Picture

Member for

1 year 3 months
Real name
SIAI Editor
Bio
SIAI Editor