Google promised its AI-generated search answers would be helpful, authoritative, and trustworthy. A growing body of evidence suggests they are frequently none of those things.
The feature known as AI Overviews — the automatically generated summaries that now appear at the top of roughly a billion daily Google searches — has been dogged by accuracy concerns since its broad rollout in May 2024. Early viral examples of bizarre errors (suggesting users put glue on pizza, or eat rocks for minerals) were initially dismissed by Google as edge cases. But reporting by The New York Times reveals that the problem runs deeper and persists longer than the company has publicly acknowledged, raising hard questions about whether the world’s dominant search engine has traded reliability for the appearance of artificial intelligence sophistication.
The core issue is deceptively simple. When a user types a query into Google, the AI Overview feature synthesizes information from across the web and presents it as a confident, authoritative paragraph — sometimes with citations, sometimes without. The format mimics the tone of an encyclopedia entry. But the underlying technology, a large language model, doesn’t actually understand truth. It predicts plausible-sounding text. And plausible isn’t the same as accurate.
According to The New York Times, independent researchers who systematically tested AI Overviews across thousands of queries found factual errors, misleading characterizations, and outright fabrications at rates that would be unacceptable for any traditional reference source. Medical queries proved particularly problematic. In some cases, AI Overviews presented outdated treatment recommendations. In others, they conflated symptoms of distinct conditions or cited studies that didn’t support the claims attributed to them.
Google has pushed back on these characterizations. The company says its internal testing shows AI Overviews are accurate at rates comparable to featured snippets, the highlighted text boxes that preceded the AI feature. Spokespeople have repeatedly pointed to the billions of queries the system handles, arguing that error rates are extremely low in percentage terms. But critics note that even a small percentage of a billion daily queries translates to millions of potentially inaccurate answers served to users every single day — users who may not have the expertise to recognize the errors.
That’s the fundamental tension. Scale.
Traditional Google search presented ten blue links and left the interpretation to the user. AI Overviews present a single synthesized answer that carries the implicit endorsement of Google itself. The shift in format is also a shift in responsibility. When a link leads to a bad source, the fault lies partly with the source. When Google’s own AI generates a wrong answer and presents it in a privileged position above all organic results, the accountability calculus changes entirely.
This matters enormously for publishers, too. The New York Times reporting highlights how AI Overviews frequently draw on journalism and expert content without driving traffic back to the original sources. Several news organizations and medical information providers told the Times that their referral traffic from Google has declined measurably since AI Overviews expanded. So the system doesn’t just risk getting things wrong — it also undermines the economic model of the very sources it depends on for training data and real-time information.
The timing of renewed scrutiny is no accident. Google is under intense competitive pressure from AI-native search products. OpenAI’s ChatGPT search feature, Perplexity AI, and Microsoft’s Copilot-integrated Bing have all made inroads with users who want conversational, synthesized answers rather than lists of links. Google’s response has been to accelerate the integration of generative AI into its core product, sometimes at the expense of caution. Internal documents referenced in earlier reporting by The Wall Street Journal showed that some Google engineers and quality raters raised concerns about the speed of AI Overview deployment, only to be overruled by leadership focused on maintaining market share.
Liz Reid, the Google executive who oversees Search, has said publicly that the company applies “the highest bar” to AI Overview quality, particularly in categories Google classifies as YMYL — Your Money or Your Life — which include health, finance, and legal information. But the independent testing described by the Times suggests this bar is not consistently met. Researchers found that YMYL queries were indeed handled more carefully on average, but errors still appeared with troubling frequency, particularly for less common medical conditions and nuanced financial questions where the model’s training data may be thinner.
One illustrative example from the Times report: a query about the safety of a specific supplement during pregnancy generated an AI Overview that cited a legitimate medical journal but mischaracterized the study’s findings, effectively reversing the authors’ conclusion. The original study found insufficient evidence to recommend the supplement. Google’s AI summary told users the supplement was “generally considered safe based on clinical research.” A subtle but dangerous inversion.
Google removed the specific answer after the Times flagged it. But the whack-a-mole nature of the correction process is itself part of the problem. With a billion queries a day generating AI Overviews, manual review of flagged errors can’t keep pace. The company relies primarily on automated quality systems and user feedback signals, which means errors can persist for days, weeks, or longer before being caught — if they’re caught at all.
The advertising implications are significant. Google’s search advertising business, which generated over $198 billion in revenue in 2024, depends on user trust. If users begin to question the reliability of AI-generated answers, their engagement patterns could shift. Early data from search analytics firms suggests that AI Overviews have already changed click-through behavior, with some queries seeing dramatically reduced clicks to organic results. That’s good for Google’s engagement metrics in the short term — users stay on Google’s page — but potentially corrosive to the trust that undergirds the entire advertising model.
Wall Street has largely given Google a pass on these concerns so far. Alphabet’s stock has performed well, buoyed by strong cloud computing growth and investor enthusiasm for AI broadly. But several analysts have begun to flag AI Overview accuracy as a longer-term risk factor. If a high-profile error causes real harm — a medical misdiagnosis acted upon, a financial decision based on fabricated data — the reputational and legal consequences could be severe.
And the legal environment is shifting. The European Union’s AI Act, which began phased enforcement in 2025, imposes transparency and accuracy requirements on AI systems that interact with consumers. While Google has argued that AI Overviews are an information retrieval tool rather than a standalone AI system subject to the Act’s strictest provisions, EU regulators have signaled they may take a broader view. In the United States, the Federal Trade Commission has opened preliminary inquiries into AI-generated content accuracy across several major platforms, though no formal enforcement actions have been announced.
There’s also the question of what happens to public knowledge over time. AI Overviews don’t just reflect the web — they shape it. When millions of users accept an AI-generated answer without clicking through to source material, that answer becomes, for practical purposes, the truth. If it’s wrong, the error propagates. Researchers at Stanford and MIT have documented what they call “epistemic closure loops” in which AI-generated content is indexed by search engines, then used as training data for the next generation of AI models, compounding initial errors across iterations. The implications for scientific literacy, public health communication, and informed democratic participation are not trivial.
Google isn’t blind to these dynamics. The company has invested heavily in what it calls “grounding” — techniques for tying AI-generated text more closely to verified source material. Recent updates to the Gemini model that powers AI Overviews have improved citation accuracy, and Google has introduced more prominent source attribution in some markets. But grounding is an incomplete solution. A model can correctly cite a source and still mischaracterize what the source says, as the pregnancy supplement example demonstrates.
Some industry observers have proposed a different approach entirely. Rather than trying to make generative AI accurate enough for high-stakes queries, Google could restrict AI Overviews to lower-risk categories — weather, basic definitions, simple factual lookups — and revert to traditional search results for anything touching health, finance, law, or other sensitive domains. This would sacrifice the wow factor of a universal AI answer but would align the technology’s capabilities with its actual reliability.
Google has shown no inclination to move in that direction. If anything, the company is expanding AI Overviews into more query categories and more countries. The competitive logic is clear: retreat from AI-generated answers and users will simply go to ChatGPT or Perplexity for the same thing, with potentially even less accuracy and accountability. It’s a race-to-the-bottom argument, and it has a certain ruthless logic. But it also means the accuracy problem is being managed rather than solved.
For the roughly 8.5 billion searches Google processes daily, the stakes are not abstract. A wrong answer about a drug interaction isn’t a fun screenshot to share on social media. It’s a potential emergency room visit. A fabricated legal precedent cited in an AI Overview isn’t a curiosity — it’s the kind of thing that has already embarrassed lawyers who relied on AI-generated briefs. The distance between a viral gaffe and genuine harm is shorter than Google’s reassurances suggest.
The company’s position, stripped to its essence, is this: AI Overviews are imperfect but improving, the error rate is low relative to volume, and the alternative — ceding the AI search market to less careful competitors — would be worse for users. Each of those claims has some merit. None of them fully addresses the core concern, which is that Google has deployed a technology at planetary scale that it cannot yet make reliably truthful, in a format that discourages the very skepticism users would need to protect themselves from its errors.
That’s not a technical problem. It’s a trust problem. And trust, once lost at scale, is extraordinarily difficult to rebuild.


WebProNews is an iEntry Publication