Google’s AI Still Can’t Count Letters in Its Own Name

Google's AI Overview confidently claims two Ps in "Google" yet often misspells the name. Recent tests from TechCrunch and Mashable show persistent failures on basic spelling and counting tasks. The root cause lies in token-based architecture that prioritizes meaning over literal characters. Despite patches and promises of fixes, the errors persist and fuel broader doubts about reliability in AI-powered search.
Google’s AI Still Can’t Count Letters in Its Own Name
Written by Eric Hastings

Google’s latest push to remake search around artificial intelligence has hit an awkward snag. Ask its AI Overview how many Ps appear in “Google” and it answers two. The system then spells the company name with an extra letter or drops one entirely. These aren’t isolated slips. They expose a stubborn weakness in the very models Google has bet its future on.

Token Limits Meet Real-World Scrutiny

Users on X noticed the pattern almost immediately after the latest updates rolled out. One post showed the AI confidently declaring exactly one “r” in the word “poop.” Another caught it claiming two “d”s in “journalism” while spelling the word j-o-u-r-n-a-d-i-s-m. Even a query about the U.S. president’s last name produced a single P but rendered it t-r-p-u-m. The mistakes spread quickly across social platforms, turning what should have been a technical footnote into public embarrassment.

But why does this keep happening? Large language models break text into tokens. These chunks represent common patterns from training data rather than individual characters. A model sees “Google” as a concept tied to search and advertising. It does not maintain a precise mental map of six distinct letters in fixed order. That gap matters. When asked to count or manipulate spelling, the system falls back on statistical likelihood instead of literal inspection.

Matthew Guzdial, an AI researcher and assistant professor at the University of Alberta, explained the mechanism in a TechCrunch interview. “When it sees the word ‘the,’ it has this one encoding of what ‘the’ means, but it does not know about ‘T,’ ‘H,’ ‘E.’” Sheridan Feucht, a PhD student studying large language model interpretability at Northeastern University, added in another TechCrunch piece that perfect tokenization remains elusive. “My guess would be that there’s no such thing as a perfect tokenizer due to this kind of fuzziness.”

Google acknowledged the specific counting problem. “Counting within words has been a known challenge for LLMs, and we’re working to fix this particular issue,” the company told TechCrunch in an emailed statement on May 27, 2026. The admission came hours after fresh examples flooded timelines. And the issue isn’t new. Two years earlier the same models stumbled on “how many r’s are in strawberry.” The pattern persists despite billions in compute and repeated fine-tuning.

Recent coverage shows the problem has not gone away. On the same day as the TechCrunch report, Mashable documented fresh tests. When asked how many e’s appear in “astronomical,” Google’s AI Overview answered two and then produced the misspelled version a-s-t-r-e-n-o-m-i-c-a-e-l. The article noted that while overall accuracy has risen, spelling tests remain a reliable way to surface errors. Users continue to share screenshots. Some express frustration. Others treat the failures as comedy.

The stakes have grown. Google placed AI Overviews at the top of search results for millions of queries. A 2026 New York Times analysis found the summaries accurate roughly nine times out of ten. That still leaves hundreds of thousands of mistakes per minute across trillions of annual searches. When those mistakes involve basic facts or simple spelling, confidence erodes. One viral tweet captured the irony: Google is revamping its entire search engine around technology that cannot reliably spell its own name.

Engineers have patched related bugs. Last week a search for “disregard” returned a strange dictionary entry that read like a chatbot deflection: “Understood. Let me know whenever you have a new prompt or question!” That flaw disappeared within days. Spelling miscues have proven harder to stamp out. Researchers say the transformer architecture itself creates the blind spot. Models predict probable next tokens. They do not parse strings the way a child learns to sound out words.

So Google finds itself in a bind. It must ship ever more ambitious AI features to stay competitive. At the same time these features invite simple tests that expose architectural limits. The company has poured resources into Gemini, the model family powering much of the new search experience. Yet the same model family still trips over letter counts that any competent spell-checker would handle without hesitation.

Industry observers point out that spelling is not the core value proposition. LLMs already generate code, summarize research papers, and draft marketing copy at speeds no human matches. But the visible failures serve as shorthand for deeper questions about reliability. If an AI cannot count letters in a six-letter word, what else might it miss in more complex domains? Users have begun to treat the outputs as first drafts rather than final answers. That shift carries consequences for publishers, advertisers, and anyone who once viewed Google as an authoritative starting point.

Google’s own history with spelling offers contrast. For years its search engine excelled at correcting user typos. The old system relied on edit distance, dictionaries, and click data. The new generative layer replaced some of that machinery with probabilistic prediction. Progress in one area created regression in another. Company statements suggest ongoing work. No timeline has emerged for a complete solution. In the meantime the jokes continue. And every new viral example reminds product teams that perception of intelligence can hinge on the most elementary tasks.

The episode also highlights tension inside the company. Teams racing to integrate AI across products must balance speed against scrutiny. Public demonstrations at events like I/O 2026 showcased agents and smarter search. Those demos rarely included prompts designed to expose token-level weaknesses. Real-world usage does exactly that. The resulting feedback loop forces iteration but also risks eroding trust in the short term.

Competitors face similar constraints. OpenAI’s models, Anthropic’s Claude, and others exhibit parallel shortcomings on letter-counting tasks. The problem appears fundamental to current scaling techniques rather than unique to Google. Still, as the dominant search provider, Google absorbs the loudest criticism. Its AI now sits atop the world’s most visited information portal. Mistakes there carry greater visibility.

Looking ahead, solutions may involve hybrid approaches. Some researchers experiment with separate spelling modules or character-level encoders that run alongside token-based models. Others explore training techniques that emphasize exact string manipulation. Whether these yield production-ready fixes before the next wave of AI search features remains uncertain. For now the public has a simple diagnostic: ask an AI to spell “Google” or count its letters. The answer often reveals more than the company intends.

Subscribe for Updates

AIDeveloper Newsletter

The AIDeveloper Email Newsletter is your essential resource for the latest in AI development. Whether you're building machine learning models or integrating AI solutions, this newsletter keeps you ahead of the curve.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us