Google’s AI Health Coach Already Hallucinates Workouts That Never Happened

Google's new AI Health Coach, tied to Fitbit data and Gemini, has begun inventing workouts and misremembering user context in early tests. Testers report phantom runs, stubborn memory errors and shallow advice despite extensive expert validation. The company acknowledges preview limitations while rolling the feature out globally. Real-world performance may determine if users trust it with their health data.
Google’s AI Health Coach Already Hallucinates Workouts That Never Happened
Written by Juan Vasquez

Early testers expected a smart companion. They got one that confidently invents exercise sessions.

Google rolled out its Health Coach feature this month, powered by Gemini and tied to the new Google Health app that replaces much of the Fitbit experience. The promise sounds straightforward. Analyze sleep logs, heart rate data, nutrition inputs and workout history. Then deliver personalized plans for fitness, recovery and daily habits. Yet within days of preview access, the system began fabricating details.

Phantom runs and stubborn memory glitches

Will Sattelberg tested the coach shortly after Google introduced the Fitbit Air tracker and the rebranded app. The AI correctly pulled his previous night’s sleep score and a legitimate workout from the day before. Then it described a five-mile run he had never taken. When challenged, the coach admitted the mistake but suggested Sattelberg simply forgot to log it. (Android Authority)

That exchange captures a core weakness. The model mixes accurate context with invented events. It does so without clear signals that its memory can fail. Adrienne So encountered similar problems during a public preview for Wired. The coach decided she was attending a work conference. She wasn’t. Later it locked onto an old statement that she felt sick and kept prescribing short, slow runs even after she corrected the record. So had to manually delete entries in the coach’s notes to reset the behavior.

Fitbit’s team acknowledged the issue in an email to So. They wrote that in the iterative public preview they expect trouble with memory expiration and persistence. Those glitches cause unexpected workout adjustments. The company said it is actively working on fixes. But the pattern raises questions about readiness for wider release.

Google launched the coach globally on May 19 for Health Premium subscribers. The service pulls data from Fitbit devices, lab results users upload, and manual logs. It offers insights on training load, sleep trends, and even summaries of medical records in some regions. Yet the company repeatedly stresses the tool is not a doctor. Responses should be checked for accuracy. Results may vary.

And they do. In one preview, the coach produced long blocks of generic encouragement that testers called shallow. The length seemed designed to project authority. The substance did not match the data at hand. So noted that relying on the coach for breakfast macros or training decisions began to change how she talked to actual people. Her husband gave her odd looks. Friends pulled back. The experience left her faster but, in her words, noticeably weirder.

These reports arrive at a difficult moment for large language models. The New York Times reported last week that newer reasoning systems from OpenAI, Google and others are producing incorrect information more often than earlier versions. Math skills have sharpened. Grasp of specific facts has slipped. Engineers cannot fully explain why. The trend matters here because health coaching demands precision. A fabricated run might seem minor. But persistent errors on recovery needs, heart rate zones or symptom interpretation carry higher stakes.

Google has invested heavily in safeguards. Its research team outlined a multi-agent architecture. One agent handles conversation. Another analyzes time-series data with code. A third brings domain expertise in fitness or sleep. The company convened a Consumer Health Advisory Panel of outside experts. It ran large user studies through Fitbit Insights Explorer and sleep labs. Most visibly, it created the SHARP framework. Safety, helpfulness, accuracy, relevance and personalization. Evaluators logged more than one million human annotations and 100,000 hours of review by specialists in cardiology, endocrinology, behavioral science and sports medicine. Autoraters now scale parts of the process. Real-world performance feeds back into updates. (Google Research Blog)

That process sounds thorough. Early results suggest gaps remain. The coach sometimes forgets context across sessions. It can blame the user for its own errors. And it generates advice that feels verbose yet surface-level. CNET reported that Google recruits external registered dietitians and internal nutrition specialists to review outputs. The company also points to its SHARP evaluations as evidence of careful work. Still, preview users keep finding hallucinations. (CNET)

Competitors have taken notice. WHOOP used the launch of Fitbit Air to highlight its own push toward human clinicians rather than pure AI guidance. Glossy examined the race to own wellness data and quoted experts who say personalization must feel supportive, not authoritative. One noted that brands need distinct experiences for different user motivations. Physical performance for some. Reassurance for others. (Glossy)

Google clearly sees the opportunity. The Health app now sits at the center of its wearable strategy. Premium access bundles with higher tiers of Gemini subscriptions. The coach can reference lab work, cycle tracking, mental wellbeing logs and more. But data privacy worries linger. So reminded readers that Fitbit information does not fuel ads yet still questioned feeding sensitive details to a system unbound by HIPAA.

Users have begun to speak up on forums. Some enjoy the accountability the coach provides for consistent running. Others report the same phantom activities and frustrating loops. One Reddit tester said the feature tracked runs well for months until it started inventing details. The pattern echoes broader AI challenges. More capable models sometimes hallucinate with greater confidence.

Google promises ongoing iteration. The public preview continues. Feedback flows directly into model tuning. The SHARP evaluations now include live performance metrics. Company spokespeople emphasize that the coach reminds users to consult professionals for medical questions. That disclaimer appears often.

Yet the early missteps matter. Health data feels personal. When an AI confidently describes workouts that never occurred, trust erodes fast. When it clings to outdated context about illness, users grow annoyed then wary. So ultimately found real running partners more motivating than her digital one. The coach helped her get faster. Human connection kept her going.

The coming weeks will test whether Google can quiet these ghosts before millions of Fitbit Air owners activate their new coach. The hardware ships May 26. The app update is already here. Hallucinations, for now, travel with it.

Subscribe for Updates

AITrends Newsletter

The AITrends Email Newsletter keeps you informed on the latest developments in artificial intelligence. Perfect for business leaders, tech professionals, and AI enthusiasts looking to stay ahead of the curve.

By signing up for our newsletter you agree to receive content related to ientry.com / webpronews.com and our affiliate partners. For additional information refer to our terms of service.

Notice an error?

Help us improve our content by reporting any issues you find.

Get the WebProNews newsletter delivered to your inbox

Get the free daily newsletter read by decision makers

Subscribe
Advertise with Us

Ready to get started?

Get our media kit

Advertise with Us