AI Chatbots Fail Most Often on Sports and Nutrition Queries, BMJ Audit Finds
A rigorous audit of five popular AI chatbots found that nearly half of their health responses were problematic, with the worst performance occurring in athletic performance and nutrition queries.
- Evidence-Based Practitioners
- Argue that AI cannot replace human expertise in sports science due to its inability to filter out commercial pseudoscience.
- Public Health Advocates
- Focus on the broader dangers of AI health misinformation and the need for regulatory oversight and user education.
- Tech Optimists
- Believe that while current models hallucinate, rapid iterations and specialized medical AI will soon bridge the accuracy gap.
Perspectives this story doesn't cover
- Registered Dietitians and Clinical Nutritionists
- Everyday Athletes Using AI
Generative artificial intelligence has rapidly become a ubiquitous tool for everyday problem-solving, with millions of users relying on chatbots to draft emails, write code, and summarize documents. In the sports world, athletes and coaches are increasingly turning to these platforms to generate marathon training plans, optimize macronutrient ratios, and seek advice on injury recovery. The appeal is obvious: chatbots deliver instant, articulate, and highly customized responses that mimic the tone of a professional consultant. However, a rigorous new audit reveals a glaring vulnerability in this technological revolution. When it comes to sports science, fluent answers are frequently masking dangerous inaccuracies.[3]
The comprehensive study, published in the peer-reviewed journal BMJ Open, was led by Dr. Nicholas B. Tiller at the Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center. The research team set out to conduct a systematic stress test of the world’s most popular public-facing AI models. They audited five major chatbots: OpenAI’s ChatGPT, Google’s Gemini, Meta AI, xAI’s Grok, and High-Flyer’s DeepSeek, evaluating their ability to handle complex, everyday health and medical queries.[1][4]
To accurately gauge the models' reliability, the researchers deployed an adversarial framework designed to mimic how real users search for information. They fed the chatbots 250 prompts across five categories known to be highly prone to online misinformation: cancer, vaccines, stem cells, nutrition, and athletic performance. The prompts included a mix of closed-ended questions, which required specific, consensus-aligned answers, and open-ended queries that invited the models to generate lists or elaborate on broader concepts.[1][5]
The top-line findings present a sobering reality check for anyone relying on AI for health guidance. According to the independent subject-matter experts who evaluated the outputs, an alarming 49.6% of all chatbot responses were rated as problematic. Breaking this down further, 30% of the answers were deemed "somewhat problematic" due to missing context or poor framing, while 19.6% were classified as "highly problematic"—meaning they contained outright misinformation or contraindicated advice that could plausibly cause harm if acted upon.[1][6]
Interestingly, the AI models did not fail uniformly across all subjects. The chatbots actually performed reasonably well when answering questions about vaccines and cancer. These are domains characterized by massive, well-structured bodies of academic research and clear scientific consensus. Because the foundational data in these fields is robust, heavily moderated, and frequently reinforced across authoritative internet sources, the models were better equipped to reproduce accurate, evidence-based content without hallucinating. The guardrails placed on these specific high-stakes medical topics by AI developers also likely contributed to their stronger performance.[1][2]
However, the chatbots collapsed when queried about athletic performance and nutrition. These two categories produced the highest frequency of problematic responses and the lowest number of accurate ones, revealing a massive blind spot in the algorithms. For athletes, practitioners, and fitness enthusiasts who view AI as a cheap, accessible alternative to a registered dietitian or a certified strength coach, this specific vulnerability is the most critical and actionable takeaway from the comprehensive audit. The models consistently failed to provide safe, reliable guidance when pushed on nuanced dietary strategies or advanced training protocols.[1][3]
The root cause of this failure lies in how large language models are trained. Sports nutrition and fitness are fields heavily saturated with commercial claims, influencer anecdotes, fad diets, and outright pseudoscience. Because chatbots pattern-match against the vast, unfiltered expanse of the internet, they struggle to separate rigorous, peer-reviewed kinesiology from aggressive supplement marketing. The models simply reflect the noise of their training data, regurgitating popular fitness myths as if they were established physiological facts.[3]
The root cause of this failure lies in how large language models are trained.
Compounding the danger of this misinformation is the unwavering certainty with which it is delivered. Despite their poor accuracy in sports science, the chatbots almost never admitted when they were out of their depth. Out of the 250 health and medical questions posed during the study, there were only two instances where a model refused to answer—both from Meta AI in response to queries about anabolic steroids and alternative cancer therapies. The rest of the time, the AI delivered flawed advice with absolute confidence.[1][4]
This confidence is often bolstered by the appearance of scientific rigor, which the study exposed as a dangerous illusion. When the researchers prompted the chatbots to provide scientific references to back up their claims, the models fell apart under scrutiny. The median completeness score for the generated citations was a dismal 40%, meaning the vast majority of the provided evidence was fundamentally flawed, incomplete, or entirely disconnected from the actual claims being made in the generated text. This lack of verifiable sourcing strips the advice of its scientific validity.[1][2]
In fact, not a single chatbot was able to produce a fully accurate reference list across the entire audit. The models frequently engaged in "citation hallucination"—inventing fake author names, generating broken DOI links, or citing papers that simply do not exist. To a layperson, these fabricated references look entirely real, creating a false veneer of scientific credibility that bypasses a user's natural skepticism.[1][3]
The audit also highlighted a pervasive issue with "false balance." When asked about controversial or unproven treatments, the chatbots often presented unscientific supplements or fad diets alongside evidence-based practices, giving them equal weight in the generated text. By failing to emphasize the overwhelming scientific consensus or the lack of clinical trials, the models inadvertently legitimize pseudoscience. This false equivalence can easily mislead an athlete looking for a competitive edge or a marginal gain, pushing them toward ineffective or potentially harmful interventions.[1][3]
The structure of the user's question also played a massive role in the likelihood of receiving bad advice. The researchers found that open-ended queries—such as "What supplements are best for overall health?"—resulted in highly problematic answers 32% of the time. In contrast, closed, yes/no questions only triggered highly problematic responses 7% of the time. This suggests that when AI is given the freedom to elaborate, it is far more likely to wander into the realm of misinformation.[1][2]
Even when the chatbots managed to provide accurate information, they often failed to communicate it effectively to a general audience. The study graded the readability of all chatbot outputs and found them to be universally classified as "Difficult." The language used was equivalent to a college sophomore-to-senior reading level, making the advice hard to digest and properly implement for the average user seeking straightforward health guidance. This high linguistic complexity further alienates users who do not have a background in sports science or medicine.[1][5]
While aggregate performance was similarly flawed across the board, there were minor differences between the models. Grok generated the highest rate of highly problematic responses, with 58% of its answers flagged overall. However, the researchers stressed that the high prevalence of problematic content reflects broader, fundamental limitations of current LLM technology, rather than the failure of any single company. No model tested is currently safe for unverified health advice.[2][4]
For the sports world, the implications of this study are clear and immediate. AI remains a powerful tool for summarizing known concepts, drafting basic schedules, or brainstorming ideas, but it simply cannot replace human expertise. Athletes, coaches, and practitioners must treat chatbots as flawed assistants rather than definitive sources of truth. To ensure safety and efficacy, users must always verify AI-generated training and nutrition plans with qualified professionals who can contextualize the advice and filter out algorithmic hallucinations.[3]
Until generative AI models are specifically siloed and trained exclusively on peer-reviewed sports science databases, they will continue to struggle with the internet's vast ocean of fitness noise. Public health advocates and researchers are now calling for stricter oversight, robust public education campaigns, and mandatory disclaimers built into the software. They warn that the uncritical adoption of these tools risks amplifying medical and nutritional misinformation on an unprecedented scale, ultimately harming the very people seeking to improve their health and performance.[1][3]
- 49.6%
- Problematic health responses
- 19.6%
- Highly problematic responses
- 40%
- Median citation completeness
- 0.8%
- Refusal to answer rate
Key points
- An audit of five major AI chatbots found that 49.6% of their responses to health queries were problematic.
- The models performed worst in the categories of athletic performance and nutrition, struggling to filter out online pseudoscience.
- Chatbots rarely refused to answer questions, delivering inaccurate advice with high confidence.
- No chatbot was able to produce a fully accurate list of scientific references, frequently hallucinating fake studies.
- Open-ended questions resulted in significantly more misinformation than closed, yes/no questions.
Sources
[1]BMJ OpenEvidence-Based PractitionersGenerative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit
Read on BMJ Open →
[2]The IndependentPublic Health AdvocatesAI chatbots often hallucinate and give inaccurate medical information, study finds
Read on The Independent →
[3]MySportScienceEvidence-Based PractitionersAI chatbots and sports nutrition: fluent answers are not always accurate
Read on MySportScience →
[4]PsyPostPublic Health AdvocatesAudit of AI chatbots reveals 49.6% of health responses are problematic
Read on PsyPost →
[5]The Straits TimesPublic Health AdvocatesAI chatbots give problematic medical advice half the time: Study
Read on The Straits Times →
[6]News-MedicalTech OptimistsStudy finds popular AI chatbots often give problematic health advice
Read on News-Medical →
Comments
More in Sports
See all →Storyline
Rowing Trades 2,000-Meter Lakes for a 500-Meter Urban Sprint in Shanghai
4 sources
Standings
Sporting CP, Kielce, and Nantes Set Blistering Pace in Expanded EHF Champions League
6 sources
Recap
U.S. Men's National Team Outlasts Canada in Five-Set Thriller to Defend NORCECA Continental Title
5 sources
Trade
Tommy Freeman Commits Long-Term Future to Northampton Saints with Contract Extension
6 sources
Every angle. Every day.
Get Sports stories with full source coverage and perspective breakdowns delivered to your inbox.




