Skip to main content
AI ReliabilityEvidence PackAug 2, 2026, 9:43 AM· 7 min read· #1 of 4 in sports

AI Chatbots Fail Most Often on Sports and Nutrition Queries, BMJ Audit Finds

A rigorous audit of five popular AI chatbots found that nearly half of their health responses were problematic, with the worst performance occurring in athletic performance and nutrition queries.

By Andres Navarro

Evidence-Based Practitioners 40%Public Health Advocates 35%Tech Optimists 25%
Evidence-Based Practitioners
Argue that AI cannot replace human expertise in sports science due to its inability to filter out commercial pseudoscience.
Public Health Advocates
Focus on the broader dangers of AI health misinformation and the need for regulatory oversight and user education.
Tech Optimists
Believe that while current models hallucinate, rapid iterations and specialized medical AI will soon bridge the accuracy gap.

Why this matters

As athletes and fitness enthusiasts increasingly rely on generative AI for training and diet plans, this research highlights that these tools confidently invent citations and struggle most with sports science, meaning users must verify AI advice with human professionals.

Key points

  • An audit of five major AI chatbots found that 49.6% of their responses to health queries were problematic.
  • The models performed worst in the categories of athletic performance and nutrition, struggling to filter out online pseudoscience.
  • Chatbots rarely refused to answer questions, delivering inaccurate advice with high confidence.
  • No chatbot was able to produce a fully accurate list of scientific references, frequently hallucinating fake studies.
  • Open-ended questions resulted in significantly more misinformation than closed, yes/no questions.
49.6%
Problematic health responses
19.6%
Highly problematic responses
40%
Median citation completeness
0.8%
Refusal to answer rate

Generative artificial intelligence has rapidly become a ubiquitous tool for everyday problem-solving, with millions of users relying on chatbots to draft emails, write code, and summarize documents. In the sports world, athletes and coaches are increasingly turning to these platforms to generate marathon training plans, optimize macronutrient ratios, and seek advice on injury recovery. The appeal is obvious: chatbots deliver instant, articulate, and highly customized responses that mimic the tone of a professional consultant. However, a rigorous new audit reveals a glaring vulnerability in this technological revolution. When it comes to sports science, fluent answers are frequently masking dangerous inaccuracies.[3]

The comprehensive study, published in the peer-reviewed journal BMJ Open, was led by Dr. Nicholas B. Tiller at the Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center. The research team set out to conduct a systematic stress test of the world’s most popular public-facing AI models. They audited five major chatbots: OpenAI’s ChatGPT, Google’s Gemini, Meta AI, xAI’s Grok, and High-Flyer’s DeepSeek, evaluating their ability to handle complex, everyday health and medical queries.[1][4]

To accurately gauge the models' reliability, the researchers deployed an adversarial framework designed to mimic how real users search for information. They fed the chatbots 250 prompts across five categories known to be highly prone to online misinformation: cancer, vaccines, stem cells, nutrition, and athletic performance. The prompts included a mix of closed-ended questions, which required specific, consensus-aligned answers, and open-ended queries that invited the models to generate lists or elaborate on broader concepts.[1][5]

The top-line findings present a sobering reality check for anyone relying on AI for health guidance. According to the independent subject-matter experts who evaluated the outputs, an alarming 49.6% of all chatbot responses were rated as problematic. Breaking this down further, 30% of the answers were deemed "somewhat problematic" due to missing context or poor framing, while 19.6% were classified as "highly problematic"—meaning they contained outright misinformation or contraindicated advice that could plausibly cause harm if acted upon.[1][6]

Nearly half of all AI chatbot responses to health and medical queries were flagged as problematic by subject-matter experts.
Nearly half of all AI chatbot responses to health and medical queries were flagged as problematic by subject-matter experts.

Interestingly, the AI models did not fail uniformly across all subjects. The chatbots actually performed reasonably well when answering questions about vaccines and cancer. These are domains characterized by massive, well-structured bodies of academic research and clear scientific consensus. Because the foundational data in these fields is robust, heavily moderated, and frequently reinforced across authoritative internet sources, the models were better equipped to reproduce accurate, evidence-based content without hallucinating. The guardrails placed on these specific high-stakes medical topics by AI developers also likely contributed to their stronger performance.[1][2]

However, the chatbots collapsed when queried about athletic performance and nutrition. These two categories produced the highest frequency of problematic responses and the lowest number of accurate ones, revealing a massive blind spot in the algorithms. For athletes, practitioners, and fitness enthusiasts who view AI as a cheap, accessible alternative to a registered dietitian or a certified strength coach, this specific vulnerability is the most critical and actionable takeaway from the comprehensive audit. The models consistently failed to provide safe, reliable guidance when pushed on nuanced dietary strategies or advanced training protocols.[1][3]

The root cause of this failure lies in how large language models are trained. Sports nutrition and fitness are fields heavily saturated with commercial claims, influencer anecdotes, fad diets, and outright pseudoscience. Because chatbots pattern-match against the vast, unfiltered expanse of the internet, they struggle to separate rigorous, peer-reviewed kinesiology from aggressive supplement marketing. The models simply reflect the noise of their training data, regurgitating popular fitness myths as if they were established physiological facts.[3]

Chatbots struggled most with sports science and nutrition, domains heavily saturated with online pseudoscience.
Chatbots struggled most with sports science and nutrition, domains heavily saturated with online pseudoscience.
The root cause of this failure lies in how large language models are trained.

Compounding the danger of this misinformation is the unwavering certainty with which it is delivered. Despite their poor accuracy in sports science, the chatbots almost never admitted when they were out of their depth. Out of the 250 health and medical questions posed during the study, there were only two instances where a model refused to answer—both from Meta AI in response to queries about anabolic steroids and alternative cancer therapies. The rest of the time, the AI delivered flawed advice with absolute confidence.[1][4]

This confidence is often bolstered by the appearance of scientific rigor, which the study exposed as a dangerous illusion. When the researchers prompted the chatbots to provide scientific references to back up their claims, the models fell apart under scrutiny. The median completeness score for the generated citations was a dismal 40%, meaning the vast majority of the provided evidence was fundamentally flawed, incomplete, or entirely disconnected from the actual claims being made in the generated text. This lack of verifiable sourcing strips the advice of its scientific validity.[1][2]

In fact, not a single chatbot was able to produce a fully accurate reference list across the entire audit. The models frequently engaged in "citation hallucination"—inventing fake author names, generating broken DOI links, or citing papers that simply do not exist. To a layperson, these fabricated references look entirely real, creating a false veneer of scientific credibility that bypasses a user's natural skepticism.[1][3]

The audit also highlighted a pervasive issue with "false balance." When asked about controversial or unproven treatments, the chatbots often presented unscientific supplements or fad diets alongside evidence-based practices, giving them equal weight in the generated text. By failing to emphasize the overwhelming scientific consensus or the lack of clinical trials, the models inadvertently legitimize pseudoscience. This false equivalence can easily mislead an athlete looking for a competitive edge or a marginal gain, pushing them toward ineffective or potentially harmful interventions.[1][3]

AI models frequently hallucinate scientific citations, creating a dangerous veneer of evidence-based credibility.
AI models frequently hallucinate scientific citations, creating a dangerous veneer of evidence-based credibility.

The structure of the user's question also played a massive role in the likelihood of receiving bad advice. The researchers found that open-ended queries—such as "What supplements are best for overall health?"—resulted in highly problematic answers 32% of the time. In contrast, closed, yes/no questions only triggered highly problematic responses 7% of the time. This suggests that when AI is given the freedom to elaborate, it is far more likely to wander into the realm of misinformation.[1][2]

Even when the chatbots managed to provide accurate information, they often failed to communicate it effectively to a general audience. The study graded the readability of all chatbot outputs and found them to be universally classified as "Difficult." The language used was equivalent to a college sophomore-to-senior reading level, making the advice hard to digest and properly implement for the average user seeking straightforward health guidance. This high linguistic complexity further alienates users who do not have a background in sports science or medicine.[1][5]

While aggregate performance was similarly flawed across the board, there were minor differences between the models. Grok generated the highest rate of highly problematic responses, with 58% of its answers flagged overall. However, the researchers stressed that the high prevalence of problematic content reflects broader, fundamental limitations of current LLM technology, rather than the failure of any single company. No model tested is currently safe for unverified health advice.[2][4]

For the sports world, the implications of this study are clear and immediate. AI remains a powerful tool for summarizing known concepts, drafting basic schedules, or brainstorming ideas, but it simply cannot replace human expertise. Athletes, coaches, and practitioners must treat chatbots as flawed assistants rather than definitive sources of truth. To ensure safety and efficacy, users must always verify AI-generated training and nutrition plans with qualified professionals who can contextualize the advice and filter out algorithmic hallucinations.[3]

Until generative AI models are specifically siloed and trained exclusively on peer-reviewed sports science databases, they will continue to struggle with the internet's vast ocean of fitness noise. Public health advocates and researchers are now calling for stricter oversight, robust public education campaigns, and mandatory disclaimers built into the software. They warn that the uncritical adoption of these tools risks amplifying medical and nutritional misinformation on an unprecedented scale, ultimately harming the very people seeking to improve their health and performance.[1][3]

How we got here

  1. 2023–2024

    Generative AI chatbots experience explosive public adoption, with millions of users beginning to use them as primary search engines for health and fitness advice.

  2. February 2025

    Researchers conduct the audit, testing ChatGPT, Gemini, Meta AI, Grok, and DeepSeek with 250 health prompts designed to stress-test their accuracy.

  3. April 2026

    The findings are published in BMJ Open, revealing that nearly half of all chatbot responses to medical and nutritional queries are problematic.

Viewpoints in depth

Sports Science Researchers

Argue that the internet's saturation with fitness pseudoscience makes general AI models inherently unreliable for athletic advice.

Researchers point out that unlike oncology or virology, sports nutrition is dominated by commercial marketing, influencer anecdotes, and conflicting dietary fads. Because large language models are trained on this vast, unfiltered web data, they naturally reproduce the noise. Experts argue that until AI tools are siloed to train exclusively on peer-reviewed kinesiology and nutrition databases, they will continue to present fad diets and unproven supplements with the same confidence as established physiological principles.

Public Health Advocates

Warn that the combination of AI's authoritative tone and fabricated citations poses a unique danger to health literacy.

Public health professionals are particularly alarmed by the 'citation hallucination' phenomenon. When a chatbot provides a confident answer backed by a fabricated journal article and a fake DOI link, it bypasses a user's natural skepticism. Advocates argue that AI companies must implement stricter guardrails, including higher refusal-to-answer rates for medical queries and mandatory disclaimers that force users to consult licensed human professionals before altering their diets or training regimens.

AI Developers & Optimists

Maintain that current flaws are temporary and that specialized medical models will soon resolve these hallucination issues.

Technology advocates acknowledge the high error rates in current public-facing models but view them as growing pains rather than permanent flaws. They emphasize that general-purpose chatbots are not designed to be medical devices. However, they point to the rapid development of specialized, medically fine-tuned models that are already achieving high accuracy on clinical exams. Optimists believe that as retrieval-augmented generation (RAG) improves, future chatbots will successfully anchor their answers to verified scientific literature, eliminating the false balance seen today.

What we don't know

  • It remains unclear how quickly AI developers can implement effective guardrails to prevent chatbots from hallucinating scientific citations.
  • Researchers do not yet know the full real-world impact of this misinformation on the actual health and training outcomes of athletes who rely on AI.
  • It is unknown whether specialized, medically fine-tuned AI models will eventually be made available to the general public for free.

Key terms

Large Language Model (LLM)
A type of artificial intelligence trained on vast amounts of text data to recognize patterns and generate human-like text responses.
Hallucination
An instance where an AI model confidently generates false, fabricated, or nonsensical information, such as inventing a fake scientific study.
False Balance
A media bias where two opposing viewpoints are presented as being equally valid, even when one is supported by overwhelming scientific consensus and the other is not.
Adversarial Prompting
A testing method where researchers deliberately ask questions designed to trick or strain an AI model into producing incorrect or harmful advice.

Frequently asked

Which AI chatbot performed the best in the study?

While performance varied slightly, aggregate accuracy did not differ significantly among the models. Grok produced the highest rate of problematic responses, but all five chatbots (including ChatGPT, Gemini, Meta AI, and DeepSeek) failed to provide consistently reliable health information.

Why did the chatbots struggle so much with sports nutrition?

Sports nutrition is a field highly prone to online misinformation, commercial claims, and fad diets. The chatbots struggled to separate rigorous, peer-reviewed sports science from the vast amount of marketing noise they were trained on.

Did the chatbots admit when they didn't know an answer?

Rarely. Out of 250 health and medical prompts, the chatbots refused to answer only twice (both by Meta AI). They consistently delivered inaccurate information with a high degree of confidence.

Can I trust the scientific references provided by AI?

No. The study found that median citation completeness was only 40%, and no chatbot produced a fully accurate reference list. Many models hallucinated authors or provided links to non-existent papers.

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Evidence-Based Practitioners 40%Public Health Advocates 35%Tech Optimists 25%
  1. [1]BMJ OpenEvidence-Based Practitioners

    Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit

    Read on BMJ Open
  2. [2]The IndependentPublic Health Advocates

    AI chatbots often hallucinate and give inaccurate medical information, study finds

    Read on The Independent
  3. [3]MySportScienceEvidence-Based Practitioners

    AI chatbots and sports nutrition: fluent answers are not always accurate

    Read on MySportScience
  4. [4]PsyPostPublic Health Advocates

    Audit of AI chatbots reveals 49.6% of health responses are problematic

    Read on PsyPost
  5. [5]The Straits TimesPublic Health Advocates

    AI chatbots give problematic medical advice half the time: Study

    Read on The Straits Times
  6. [6]News-MedicalTech Optimists

    Study finds popular AI chatbots often give problematic health advice

    Read on News-Medical
Stay informed

Every angle. Every day.

Get sports stories with full source coverage and perspective breakdowns delivered to your inbox.