AI Chatbots Fail Most Often on Sports and Nutrition Queries, BMJ Audit Finds
A rigorous audit of five popular AI chatbots found that nearly half of their health responses were problematic, with the worst performance occurring in athletic performance and nutrition queries.
- Evidence-Based Practitioners
- Argue that AI cannot replace human expertise in sports science due to its inability to filter out commercial pseudoscience.
- Public Health Advocates
- Focus on the broader dangers of AI health misinformation and the need for regulatory oversight and user education.
- Tech Optimists
- Believe that while current models hallucinate, rapid iterations and specialized medical AI will soon bridge the accuracy gap.
Why this matters
As athletes and fitness enthusiasts increasingly rely on generative AI for training and diet plans, this research highlights that these tools confidently invent citations and struggle most with sports science, meaning users must verify AI advice with human professionals.
Key points
- An audit of five major AI chatbots found that 49.6% of their responses to health queries were problematic.
- The models performed worst in the categories of athletic performance and nutrition, struggling to filter out online pseudoscience.
- Chatbots rarely refused to answer questions, delivering inaccurate advice with high confidence.
- No chatbot was able to produce a fully accurate list of scientific references, frequently hallucinating fake studies.
- Open-ended questions resulted in significantly more misinformation than closed, yes/no questions.
Generative artificial intelligence has rapidly become a ubiquitous tool for everyday problem-solving, with millions of users relying on chatbots to draft emails, write code, and summarize documents. In the sports world, athletes and coaches are increasingly turning to these platforms to generate marathon training plans, optimize macronutrient ratios, and seek advice on injury recovery. The appeal is obvious: chatbots deliver instant, articulate, and highly customized responses that mimic the tone of a professional consultant. However, a rigorous new audit reveals a glaring vulnerability in this technological revolution. When it comes to sports science, fluent answers are frequently masking dangerous inaccuracies.[3]
The comprehensive study, published in the peer-reviewed journal BMJ Open, was led by Dr. Nicholas B. Tiller at the Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center. The research team set out to conduct a systematic stress test of the world’s most popular public-facing AI models. They audited five major chatbots: OpenAI’s ChatGPT, Google’s Gemini, Meta AI, xAI’s Grok, and High-Flyer’s DeepSeek, evaluating their ability to handle complex, everyday health and medical queries.[1][4]
To accurately gauge the models' reliability, the researchers deployed an adversarial framework designed to mimic how real users search for information. They fed the chatbots 250 prompts across five categories known to be highly prone to online misinformation: cancer, vaccines, stem cells, nutrition, and athletic performance. The prompts included a mix of closed-ended questions, which required specific, consensus-aligned answers, and open-ended queries that invited the models to generate lists or elaborate on broader concepts.[1][5]
The top-line findings present a sobering reality check for anyone relying on AI for health guidance. According to the independent subject-matter experts who evaluated the outputs, an alarming 49.6% of all chatbot responses were rated as problematic. Breaking this down further, 30% of the answers were deemed "somewhat problematic" due to missing context or poor framing, while 19.6% were classified as "highly problematic"—meaning they contained outright misinformation or contraindicated advice that could plausibly cause harm if acted upon.[1][6]

Interestingly, the AI models did not fail uniformly across all subjects. The chatbots actually performed reasonably well when answering questions about vaccines and cancer. These are domains characterized by massive, well-structured bodies of academic research and clear scientific consensus. Because the foundational data in these fields is robust, heavily moderated, and frequently reinforced across authoritative internet sources, the models were better equipped to reproduce accurate, evidence-based content without hallucinating. The guardrails placed on these specific high-stakes medical topics by AI developers also likely contributed to their stronger performance.[1][2]
However, the chatbots collapsed when queried about athletic performance and nutrition. These two categories produced the highest frequency of problematic responses and the lowest number of accurate ones, revealing a massive blind spot in the algorithms. For athletes, practitioners, and fitness enthusiasts who view AI as a cheap, accessible alternative to a registered dietitian or a certified strength coach, this specific vulnerability is the most critical and actionable takeaway from the comprehensive audit. The models consistently failed to provide safe, reliable guidance when pushed on nuanced dietary strategies or advanced training protocols.[1][3]
The root cause of this failure lies in how large language models are trained. Sports nutrition and fitness are fields heavily saturated with commercial claims, influencer anecdotes, fad diets, and outright pseudoscience. Because chatbots pattern-match against the vast, unfiltered expanse of the internet, they struggle to separate rigorous, peer-reviewed kinesiology from aggressive supplement marketing. The models simply reflect the noise of their training data, regurgitating popular fitness myths as if they were established physiological facts.[3]

The root cause of this failure lies in how large language models are trained.
Compounding the danger of this misinformation is the unwavering certainty with which it is delivered. Despite their poor accuracy in sports science, the chatbots almost never admitted when they were out of their depth. Out of the 250 health and medical questions posed during the study, there were only two instances where a model refused to answer—both from Meta AI in response to queries about anabolic steroids and alternative cancer therapies. The rest of the time, the AI delivered flawed advice with absolute confidence.[1][4]
This confidence is often bolstered by the appearance of scientific rigor, which the study exposed as a dangerous illusion. When the researchers prompted the chatbots to provide scientific references to back up their claims, the models fell apart under scrutiny. The median completeness score for the generated citations was a dismal 40%, meaning the vast majority of the provided evidence was fundamentally flawed, incomplete, or entirely disconnected from the actual claims being made in the generated text. This lack of verifiable sourcing strips the advice of its scientific validity.[1][2]
In fact, not a single chatbot was able to produce a fully accurate reference list across the entire audit. The models frequently engaged in "citation hallucination"—inventing fake author names, generating broken DOI links, or citing papers that simply do not exist. To a layperson, these fabricated references look entirely real, creating a false veneer of scientific credibility that bypasses a user's natural skepticism.[1][3]
The audit also highlighted a pervasive issue with "false balance." When asked about controversial or unproven treatments, the chatbots often presented unscientific supplements or fad diets alongside evidence-based practices, giving them equal weight in the generated text. By failing to emphasize the overwhelming scientific consensus or the lack of clinical trials, the models inadvertently legitimize pseudoscience. This false equivalence can easily mislead an athlete looking for a competitive edge or a marginal gain, pushing them toward ineffective or potentially harmful interventions.[1][3]

The structure of the user's question also played a massive role in the likelihood of receiving bad advice. The researchers found that open-ended queries—such as "What supplements are best for overall health?"—resulted in highly problematic answers 32% of the time. In contrast, closed, yes/no questions only triggered highly problematic responses 7% of the time. This suggests that when AI is given the freedom to elaborate, it is far more likely to wander into the realm of misinformation.[1][2]
Even when the chatbots managed to provide accurate information, they often failed to communicate it effectively to a general audience. The study graded the readability of all chatbot outputs and found them to be universally classified as "Difficult." The language used was equivalent to a college sophomore-to-senior reading level, making the advice hard to digest and properly implement for the average user seeking straightforward health guidance. This high linguistic complexity further alienates users who do not have a background in sports science or medicine.[1][5]
While aggregate performance was similarly flawed across the board, there were minor differences between the models. Grok generated the highest rate of highly problematic responses, with 58% of its answers flagged overall. However, the researchers stressed that the high prevalence of problematic content reflects broader, fundamental limitations of current LLM technology, rather than the failure of any single company. No model tested is currently safe for unverified health advice.[2][4]
For the sports world, the implications of this study are clear and immediate. AI remains a powerful tool for summarizing known concepts, drafting basic schedules, or brainstorming ideas, but it simply cannot replace human expertise. Athletes, coaches, and practitioners must treat chatbots as flawed assistants rather than definitive sources of truth. To ensure safety and efficacy, users must always verify AI-generated training and nutrition plans with qualified professionals who can contextualize the advice and filter out algorithmic hallucinations.[3]
Until generative AI models are specifically siloed and trained exclusively on peer-reviewed sports science databases, they will continue to struggle with the internet's vast ocean of fitness noise. Public health advocates and researchers are now calling for stricter oversight, robust public education campaigns, and mandatory disclaimers built into the software. They warn that the uncritical adoption of these tools risks amplifying medical and nutritional misinformation on an unprecedented scale, ultimately harming the very people seeking to improve their health and performance.[1][3]
How we got here
2023–2024
Generative AI chatbots experience explosive public adoption, with millions of users beginning to use them as primary search engines for health and fitness advice.
February 2025
Researchers conduct the audit, testing ChatGPT, Gemini, Meta AI, Grok, and DeepSeek with 250 health prompts designed to stress-test their accuracy.
April 2026
The findings are published in BMJ Open, revealing that nearly half of all chatbot responses to medical and nutritional queries are problematic.
Viewpoints in depth
Sports Science Researchers
Argue that the internet's saturation with fitness pseudoscience makes general AI models inherently unreliable for athletic advice.
Researchers point out that unlike oncology or virology, sports nutrition is dominated by commercial marketing, influencer anecdotes, and conflicting dietary fads. Because large language models are trained on this vast, unfiltered web data, they naturally reproduce the noise. Experts argue that until AI tools are siloed to train exclusively on peer-reviewed kinesiology and nutrition databases, they will continue to present fad diets and unproven supplements with the same confidence as established physiological principles.
Public Health Advocates
Warn that the combination of AI's authoritative tone and fabricated citations poses a unique danger to health literacy.
Public health professionals are particularly alarmed by the 'citation hallucination' phenomenon. When a chatbot provides a confident answer backed by a fabricated journal article and a fake DOI link, it bypasses a user's natural skepticism. Advocates argue that AI companies must implement stricter guardrails, including higher refusal-to-answer rates for medical queries and mandatory disclaimers that force users to consult licensed human professionals before altering their diets or training regimens.
AI Developers & Optimists
Maintain that current flaws are temporary and that specialized medical models will soon resolve these hallucination issues.
Technology advocates acknowledge the high error rates in current public-facing models but view them as growing pains rather than permanent flaws. They emphasize that general-purpose chatbots are not designed to be medical devices. However, they point to the rapid development of specialized, medically fine-tuned models that are already achieving high accuracy on clinical exams. Optimists believe that as retrieval-augmented generation (RAG) improves, future chatbots will successfully anchor their answers to verified scientific literature, eliminating the false balance seen today.
What we don't know
- It remains unclear how quickly AI developers can implement effective guardrails to prevent chatbots from hallucinating scientific citations.
- Researchers do not yet know the full real-world impact of this misinformation on the actual health and training outcomes of athletes who rely on AI.
- It is unknown whether specialized, medically fine-tuned AI models will eventually be made available to the general public for free.
Key terms
- Large Language Model (LLM)
- A type of artificial intelligence trained on vast amounts of text data to recognize patterns and generate human-like text responses.
- Hallucination
- An instance where an AI model confidently generates false, fabricated, or nonsensical information, such as inventing a fake scientific study.
- False Balance
- A media bias where two opposing viewpoints are presented as being equally valid, even when one is supported by overwhelming scientific consensus and the other is not.
- Adversarial Prompting
- A testing method where researchers deliberately ask questions designed to trick or strain an AI model into producing incorrect or harmful advice.
Frequently asked
Which AI chatbot performed the best in the study?
While performance varied slightly, aggregate accuracy did not differ significantly among the models. Grok produced the highest rate of problematic responses, but all five chatbots (including ChatGPT, Gemini, Meta AI, and DeepSeek) failed to provide consistently reliable health information.
Why did the chatbots struggle so much with sports nutrition?
Sports nutrition is a field highly prone to online misinformation, commercial claims, and fad diets. The chatbots struggled to separate rigorous, peer-reviewed sports science from the vast amount of marketing noise they were trained on.
Did the chatbots admit when they didn't know an answer?
Rarely. Out of 250 health and medical prompts, the chatbots refused to answer only twice (both by Meta AI). They consistently delivered inaccurate information with a high degree of confidence.
Can I trust the scientific references provided by AI?
No. The study found that median citation completeness was only 40%, and no chatbot produced a fully accurate reference list. Many models hallucinated authors or provided links to non-existent papers.
Sources
[1]BMJ OpenEvidence-Based Practitioners
Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit
Read on BMJ Open →[2]The IndependentPublic Health Advocates
AI chatbots often hallucinate and give inaccurate medical information, study finds
Read on The Independent →[3]MySportScienceEvidence-Based Practitioners
AI chatbots and sports nutrition: fluent answers are not always accurate
Read on MySportScience →[4]PsyPostPublic Health Advocates
Audit of AI chatbots reveals 49.6% of health responses are problematic
Read on PsyPost →[5]The Straits TimesPublic Health Advocates
AI chatbots give problematic medical advice half the time: Study
Read on The Straits Times →[6]News-MedicalTech Optimists
Study finds popular AI chatbots often give problematic health advice
Read on News-Medical →
Every angle. Every day.
Get sports stories with full source coverage and perspective breakdowns delivered to your inbox.









