Early testers expected a smart companion. They got one that confidently invents exercise sessions.
Google rolled out its Health Coach feature this month, powered by Gemini and tied to the new Google Health app that replaces much of the Fitbit experience. The promise sounds straightforward. Analyze sleep logs, heart rate data, nutrition inputs and workout history. Then deliver personalized plans for fitness, recovery and daily habits. Yet within days of preview access, the system began fabricating details.
Phantom runs and stubborn memory glitches
Will Sattelberg tested the coach shortly after Google introduced the Fitbit Air tracker and the rebranded app. The AI correctly pulled his previous night’s sleep score and a legitimate workout from the day before. Then it described a five-mile run he had never taken. When challenged, the coach admitted the mistake but suggested Sattelberg simply forgot to log it. (Android Authority)
That exchange captures a core weakness. The model mixes accurate context with invented events. It does so without clear signals that its memory can fail. Adrienne So encountered similar problems during a public preview for Wired. The coach decided she was attending a work conference. She wasn’t. Later it locked onto an old statement that she felt sick and kept prescribing short, slow runs even after she corrected the record. So had to manually delete entries in the coach’s notes to reset the behavior.
Fitbit’s team acknowledged the issue in an email to So. They wrote that in the iterative public preview they expect trouble with memory expiration and persistence. Those glitches cause unexpected workout adjustments. The company said it is actively working on fixes. But the pattern raises questions about readiness for wider release.
Google launched the coach globally on May 19 for Health Premium subscribers. The service pulls data from Fitbit devices, lab results users upload, and manual logs. It offers insights on training load, sleep trends, and even summaries of medical records in some regions. Yet the company repeatedly stresses the tool is not a doctor. Responses should be checked for accuracy. Results may vary.
And they do. In one preview, the coach produced long blocks of generic encouragement that testers called shallow. The length seemed designed to project authority. The substance did not match the data at hand. So noted that relying on the coach for breakfast macros or training decisions began to change how she talked to actual people. Her husband gave her odd looks. Friends pulled back. The experience left her faster but, in her words, noticeably weirder.
These reports arrive at a difficult moment for large language models. The New York Times reported last week that newer reasoning systems from OpenAI, Google and others are producing incorrect information more often than earlier versions. Math skills have sharpened. Grasp of specific facts has slipped. Engineers cannot fully explain why. The trend matters here because health coaching demands precision. A fabricated run might seem minor. But persistent errors on recovery needs, heart rate zones or symptom interpretation carry higher stakes.
Google has invested heavily in safeguards. Its research team outlined a multi-agent architecture. One agent handles conversation. Another analyzes time-series data with code. A third brings domain expertise in fitness or sleep. The company convened a Consumer Health Advisory Panel of outside experts. It ran large user studies through Fitbit Insights Explorer and sleep labs. Most visibly, it created the SHARP framework. Safety, helpfulness, accuracy, relevance and personalization. Evaluators logged more than one million human annotations and 100,000 hours of review by specialists in cardiology, endocrinology, behavioral science and sports medicine. Autoraters now scale parts of the process. Real-world performance feeds back into updates. (Google Research Blog)
That process sounds thorough. Early results suggest gaps remain. The coach sometimes forgets context across sessions. It can blame the user for its own errors. And it generates advice that feels verbose yet surface-level. CNET reported that Google recruits external registered dietitians and internal nutrition specialists to review outputs. The company also points to its SHARP evaluations as evidence of careful work. Still, preview users keep finding hallucinations. (CNET)
Competitors have taken notice. WHOOP used the launch of Fitbit Air to highlight its own push toward human clinicians rather than pure AI guidance. Glossy examined the race to own wellness data and quoted experts who say personalization must feel supportive, not authoritative. One noted that brands need distinct experiences for different user motivations. Physical performance for some. Reassurance for others. (Glossy)
Google clearly sees the opportunity. The Health app now sits at the center of its wearable strategy. Premium access bundles with higher tiers of Gemini subscriptions. The coach can reference lab work, cycle tracking, mental wellbeing logs and more. But data privacy worries linger. So reminded readers that Fitbit information does not fuel ads yet still questioned feeding sensitive details to a system unbound by HIPAA.
Users have begun to speak up on forums. Some enjoy the accountability the coach provides for consistent running. Others report the same phantom activities and frustrating loops. One Reddit tester said the feature tracked runs well for months until it started inventing details. The pattern echoes broader AI challenges. More capable models sometimes hallucinate with greater confidence.
Google promises ongoing iteration. The public preview continues. Feedback flows directly into model tuning. The SHARP evaluations now include live performance metrics. Company spokespeople emphasize that the coach reminds users to consult professionals for medical questions. That disclaimer appears often.
Yet the early missteps matter. Health data feels personal. When an AI confidently describes workouts that never occurred, trust erodes fast. When it clings to outdated context about illness, users grow annoyed then wary. So ultimately found real running partners more motivating than her digital one. The coach helped her get faster. Human connection kept her going.
The coming weeks will test whether Google can quiet these ghosts before millions of Fitbit Air owners activate their new coach. The hardware ships May 26. The app update is already here. Hallucinations, for now, travel with it.


WebProNews is an iEntry Publication