The Machines Are Reading Everything: How AI Bot Traffic Is Overwhelming Publishers and Rewriting the Rules of the Web

AI bot traffic has surged dramatically in 2025, with some publishers reporting automated crawlers now exceed human visitors. The resulting infrastructure costs, content extraction without compensation, and declining referral traffic are creating an existential crisis for digital media.
The Machines Are Reading Everything: How AI Bot Traffic Is Overwhelming Publishers and Rewriting the Rules of the Web
Written by Sara Donnelly

The bots won’t stop coming.

Over the past year, a massive and accelerating wave of AI-driven web crawlers has descended on publishers’ websites, consuming server resources, scraping content at industrial scale, and in many cases doing so without permission, compensation, or even basic courtesy. The traffic isn’t marginal. According to new data reported by Search Engine Land, AI bot traffic surged dramatically in early 2025, with some publishers reporting that automated crawlers now account for a staggering share of their total server requests — in certain cases exceeding human traffic entirely.

This isn’t a theoretical problem anymore. It’s an operational crisis.

The numbers tell a story that should alarm anyone who produces original content on the internet. Data compiled from server logs, analytics platforms, and publisher reports show that AI crawlers from companies like OpenAI, Anthropic, Google, Meta, Apple, Amazon, and a growing constellation of smaller AI startups are hitting websites millions of times per day. These bots are pulling text, images, and structured data to feed the insatiable training pipelines of large language models and to populate the AI-generated answers that are increasingly replacing traditional search results. Some crawlers identify themselves honestly. Many don’t. And the volume has exploded — up by multiples, not percentages, compared to a year ago.

The scale is hard to overstate. Condé Nast, one of the world’s largest magazine publishers, disclosed earlier this year that it was seeing roughly 100 million AI bot requests per day across its properties, which include The New Yorker, Wired, Vogue, and Vanity Fair. That figure represented a sharp increase from previous periods and was large enough to materially affect infrastructure costs. Other major publishers have shared similar stories privately, describing a kind of digital siege in which their servers are pounded around the clock by automated agents that extract value while contributing nothing to the publishers’ bottom lines.

The economics are brutal. Every bot request costs money — bandwidth, compute, CDN fees. But those bots don’t see ads. They don’t subscribe. They don’t click. They take content and leave.

What makes the current moment particularly fraught is the sheer number of actors involved. It was one thing when Googlebot was the dominant crawler and publishers understood the implicit bargain: let Google index your content, and in return you get search traffic. That deal, imperfect as it was, at least had a logic to it. Now there are dozens of AI crawlers, each with its own user-agent string — or no identifiable string at all — and the bargain has collapsed. OpenAI’s GPTBot, Anthropic’s ClaudeBot, ByteDance’s Bytespider, Apple’s Applebot-Extended, and numerous others are all requesting the same pages, often within minutes of each other. The result is a multiplication of load with no corresponding multiplication of benefit.

Publishers have tried to fight back. The standard tool has been robots.txt, the decades-old protocol that allows website owners to signal which crawlers are welcome and which are not. But the protocol is voluntary, and compliance has been inconsistent at best. Search Engine Land reported that many publishers who blocked specific AI bots in their robots.txt files saw little or no reduction in unwanted crawl traffic, suggesting that some bots either ignore the directive or operate under unrecognized user-agent strings. It’s the digital equivalent of posting a “No Trespassing” sign that trespassers can’t read — or choose not to.

Some publishers have gone further. The New York Times filed a high-profile lawsuit against OpenAI and Microsoft in late 2023, alleging that the companies used its copyrighted journalism to train AI models without authorization. That case is still working through the courts. Other outlets, including the Chicago Tribune‘s parent company and several Alden Global Capital-owned newspapers, have filed similar suits. But litigation is slow, expensive, and uncertain. And it doesn’t solve the immediate technical problem of servers buckling under bot load.

The financial dimension deserves close scrutiny. Major AI companies are spending tens of billions of dollars building and training models. OpenAI’s annualized revenue reportedly exceeded $5 billion in early 2025. Anthropic has raised over $10 billion in funding. Google’s parent Alphabet generated $350 billion in revenue last year. These are among the most valuable and well-capitalized companies on Earth. And yet the content that makes their AI products useful — the journalism, the analysis, the creative writing, the reference material — is being acquired for free, or close to it, from publishers whose business models are already under severe strain.

A few licensing deals have been struck. OpenAI has signed agreements with the Associated Press, Axel Springer, Le Monde, Prisa Media, and several other outlets. Google has its own set of publisher partnerships. But the total dollars flowing to publishers through these arrangements are modest relative to the value being extracted. And smaller publishers — local newspapers, niche trade publications, independent blogs — have no leverage to negotiate such deals and no practical way to prevent their content from being ingested.

The technical arms race is intensifying. Some publishers have deployed bot-detection services from companies like Cloudflare, Akamai, and Vercel to identify and throttle AI crawlers. Cloudflare in particular has rolled out features specifically designed to block AI bots, and its data has provided some of the most granular public information about crawl patterns. According to Cloudflare’s analysis, AI bot traffic increased substantially across its network in the first quarter of 2025, with certain sectors — news, reference, and e-commerce — bearing the heaviest load.

But blocking bots creates its own problems. Publishers worry that blocking Google’s AI crawlers might hurt their visibility in Google’s AI Overviews, the AI-generated summaries that now appear at the top of many search results pages. It’s a Catch-22: allow the crawling and lose direct traffic as users get answers without clicking through, or block the crawling and risk disappearing from AI-powered search results altogether. Neither option is good.

And then there’s the traffic cliff. Multiple studies, including research from analytics firms Datos and Similarweb, have shown that AI Overviews and AI chatbots are reducing click-through rates to publisher websites. When a user asks ChatGPT or Google’s Gemini a question and gets a synthesized answer drawn from multiple sources, the incentive to visit the original source evaporates. Some publishers have reported referral traffic declines of 20% to 40% from Google searches where AI Overviews appear. That’s not a rounding error. That’s an existential threat for ad-supported digital media.

The situation is creating strange bedfellows. Publishers that have competed fiercely for decades are now finding common cause in lobbying efforts aimed at Congress and regulatory agencies. The News/Media Alliance, a trade group representing thousands of publishers, has been pushing for federal legislation that would give news organizations the right to collectively negotiate with AI companies — essentially an antitrust exemption similar to what Australia’s News Media Bargaining Code provided. The Journalism Competition and Preservation Act, which has been introduced in various forms over the past few years, would do exactly this. But the bill has stalled repeatedly, caught in the broader political crosswinds around AI regulation and tech policy.

In Europe, the regulatory environment is more aggressive. The EU’s AI Act, which began taking effect in stages in 2024 and 2025, includes transparency requirements for AI training data. The EU Copyright Directive gives publishers certain rights over the use of their content by online platforms. Several European publishers have filed complaints with national regulators, and France’s competition authority has already taken action against Google over related issues. But enforcement is uneven, and the global nature of AI training makes jurisdictional boundaries somewhat academic.

Meanwhile, the AI companies themselves are sending mixed signals. OpenAI has publicly stated that it respects robots.txt and offers a mechanism for publishers to opt out of training data collection. But reports from publishers suggest the reality is more complicated. Some have documented crawl activity from IP addresses associated with AI companies even after implementing blocks. Others have noted that content they explicitly excluded from AI training still appears, paraphrased or summarized, in AI model outputs. The provenance problem — proving that a specific piece of content was used in training — remains one of the thorniest technical and legal challenges in the entire debate.

Smaller AI companies and startups are even less restrained. Many operate crawlers that don’t identify themselves as AI-related, making them difficult to distinguish from legitimate search engine bots or regular user traffic. Some use residential proxy networks to mask their origin. The result is a kind of dark web of content acquisition, invisible to most publishers and nearly impossible to police.

So where does this go? The optimistic scenario involves a market-based solution: AI companies recognize that high-quality content is a competitive advantage, pay fair licensing fees, and publishers develop new revenue streams from AI partnerships. Some version of this is already happening at the top of the market, with major outlets negotiating deals worth tens of millions of dollars. But the total addressable market for such deals is limited, and the vast majority of publishers will never be large enough to attract a licensing offer.

The pessimistic scenario is darker. AI companies continue to scrape freely, publishers lose traffic and revenue, journalism contracts further, and the quality of information available on the open web degrades — which in turn degrades the quality of AI training data, creating a vicious cycle that some researchers have called “model collapse.” If AI models are increasingly trained on AI-generated content rather than original human work, the outputs become progressively less reliable, less original, and less useful. The web eats itself.

The most likely outcome falls somewhere between these poles. Regulation will come, but slowly and imperfectly. Licensing deals will expand, but unevenly. Technical countermeasures will improve, but so will the bots. And publishers will continue to bear a disproportionate share of the cost of an AI boom that depends fundamentally on their work.

One thing is clear. The implicit social contract of the open web — the idea that content creators publish freely and benefit from the traffic and attention that aggregators and search engines send their way — is breaking down. AI companies are extracting value at a scale and speed that the old model was never designed to handle. The question isn’t whether the model needs to change. It already has. The question is who gets to write the new rules, and whether publishers will have any meaningful say in the process.

For now, the bots keep crawling. And the meter keeps running.

Subscribe for Updates

AITrends Newsletter

The AITrends Email Newsletter keeps you informed on the latest developments in artificial intelligence. Perfect for business leaders, tech professionals, and AI enthusiasts looking to stay ahead of the curve.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us