Research IntegrityEvidence PackJul 9, 2026, 11:20 AM· 5 min read· #5 of 5 in ai

AI-Generated False Citations Found in Official Government Reports, Undermining Policy Integrity

Three official policy documents from the White House, the European Union, and South Africa were found to contain fabricated academic citations generated by artificial intelligence. The discoveries have prompted policy experts to classify AI-assisted research hallucinations as a critical vulnerability to national security and regulatory integrity.

By Factlen Editorial Team

Policy Researchers & Academics 40%Government Agencies 30%AI Developers 15%Journalists & Watchdogs 15%
Policy Researchers & Academics
Focus on research integrity, mandatory verification, and the contamination of the scholarly record.
Government Agencies
Focus on the operational need for speed versus the embarrassment of retractions, implementing new internal review processes.
AI Developers
Focus on improving model grounding, Retrieval-Augmented Generation, and the limitations of current statistical inference.
Journalists & Watchdogs
Focus on exposing institutional vulnerabilities and holding governments accountable for AI misuse.

What's not represented

  • · Citizens affected by policies based on fabricated data
  • · Traditional academic publishers managing the influx of AI-generated research

Why this matters

When government agencies rely on fabricated AI-generated research to draft policy, the foundation of federal budgets, health regulations, and national security assessments is compromised. This vulnerability exposes how the unverified use of artificial intelligence can quietly corrupt the information supply chain that dictates real-world laws and operational decisions.

Key points

  • Three official government documents from the US, EU, and South Africa were found to contain fabricated AI-generated citations.
  • The White House health assessment included 'oaicite' markers, revealing the use of OpenAI's ChatGPT in drafting the report.
  • A 2023 study found that 55% of GPT-3.5 citations and 18% of GPT-4 citations were entirely fabricated.
  • Policy experts warn that AI hallucinations in government reports pose a severe threat to national security and regulatory integrity.
  • Researchers are calling for automated citation verification to become mandatory infrastructure in the government publishing pipeline.
26
Incorrect footnotes in ENISA report
55%
Fabricated citations by GPT-3.5 (2023)
18%
Fabricated citations by GPT-4 (2023)
6
Fake citations in withdrawn SA policy

The transition of AI hallucinations from an academic embarrassment to a systemic policy vulnerability is now complete. Between May 2025 and April 2026, three separate official government documents—spanning a European regulator, a national government, and the White House—were published containing fabricated academic citations generated by artificial intelligence.[1]

The discovery highlights a critical institutional blind spot. While governments worldwide have heavily invested in securing AI models against cyberattacks and catastrophic risks, they have largely overlooked the subtle corruption of the information supply chain. When policy is built on non-existent research, the foundation of regulatory and operational decisions begins to crack.[1][4]

The most prominent incident occurred within the United States government. A flagship "Make Our Children Healthy Assessment" released by the White House cited multiple scientific studies that simply did not exist, while misattributing findings to genuine authors who had never conducted the referenced research.[2]

The fabrication was first flagged by the nonprofit news publication NOTUS, which identified incorrect formatting, missing authors, and mismatched issue numbers in the report's bibliography. Following the inquiry, the report was quietly pulled, scrubbed of the fake citations, and reissued with authentic sources.

Three high-profile instances of AI hallucinations infiltrating official government policy documents over the past year.
Three high-profile instances of AI hallucinations infiltrating official government policy documents over the past year.

Further investigation by The Washington Post revealed the mechanical cause of the error. Journalists found dozens of footnotes in the original document containing an "oaicite" marker—a distinct digital fingerprint left behind by OpenAI's ChatGPT when it generates text and attempts to format references.[2]

Across the Atlantic, the European Union Agency for Cybersecurity (ENISA) faced a similar breach of research integrity. The agency's widely read threat landscape report was found by outside researchers to contain AI-generated or otherwise erroneous references that pointed to non-existent cybersecurity frameworks.[1]

According to an analysis detailed in TechPolicy.Press, one ENISA report contained 26 incorrect footnotes out of a total of 492. The cybersecurity agency was forced to issue a revised version of the document to correct the fabricated links, an embarrassing misstep for an organization dedicated to digital security and threat mitigation.[1]

According to an analysis detailed in TechPolicy.Press, one ENISA report contained 26 incorrect footnotes out of a total of 492.

In the Global South, the consequences of AI-assisted drafting were even more disruptive. The South African government was forced to completely withdraw its draft national AI policy after media reporting revealed that at least six of its sixty-seven academic citations were entirely fabricated.[1]

The South African government was forced to withdraw its draft national AI policy after fabricated citations were discovered.
The South African government was forced to withdraw its draft national AI policy after fabricated citations were discovered.

Rachel A. George, a Stanford lecturer and nonresident fellow at the Carnegie Endowment for International Peace, has cataloged these failures, arguing that research integrity must now be treated as a national security problem. Threat assessments and health policies directly dictate federal budgets, regulatory frameworks, and operational deployments.[1][4]

The underlying mechanism driving these fabrications is a well-documented flaw in large language models (LLMs) known as hallucination. When an AI model cannot find an exact match for a citation within its training data, it does not simply return a blank result. Instead, it relies on statistical inference to generate a plausible-sounding reference, complete with real author names, realistic journal titles, and standard formatting.[3]

Because these models are optimized for fluency and coherence rather than strict factual retrieval, they excel at mimicking the cadence of academic writing. The model cannot distinguish between a real citation it encountered in a legitimate academic paper and a fake citation it generated based on statistical probabilities.

A 2023 study published in Scientific Reports quantified the scale of this problem. Researchers asking ChatGPT to generate literature reviews found that 55 percent of citations produced by GPT-3.5, and 18 percent of those from the more advanced GPT-4, were entirely fabricated. More recent evaluations suggest the problem persists; a 2025 analysis of OpenAI's GPT-4o found that the model generated fabricated citations roughly 20 percent of the time.[1][3]

Studies show that even advanced AI models continue to fabricate a significant percentage of academic citations.
Studies show that even advanced AI models continue to fabricate a significant percentage of academic citations.

Crucially, the rate of AI hallucination increases significantly when models are queried about specialized, niche, or highly technical topics. This is precisely the domain where government policy analysis operates, making federal agencies uniquely vulnerable to AI-generated misinformation.[1]

The integration of web-browsing capabilities and Model Context Protocols (MCPs) into AI tools was supposed to anchor models to live data and reduce these hallucinations. However, experts note that models still struggle to distinguish between a legitimate academic paper and a fabricated citation encountered in a low-quality blog post.[1][3]

To combat this vulnerability, policy experts are advocating for a fundamental shift in how governments handle document publication. Researchers argue that automated citation verification must become standard infrastructure in the publishing pipeline, rather than an optional final check left to overburdened human staffers.[1]

The mechanism of an AI hallucination: how statistical inference creates plausible but entirely fake academic references.
The mechanism of an AI hallucination: how statistical inference creates plausible but entirely fake academic references.

Furthermore, there is a growing push for mandatory disclosure standards whenever AI is used in the drafting of government research. Just as financial conflicts of interest must be declared, the use of generative models in synthesizing policy data may soon require explicit attribution on the cover page of federal reports.[1][4]

Until these safeguards are institutionalized, the burden falls entirely on human reviewers. Agencies are being urged to invest in specialized training for staff to identify the subtle "tells" of AI-generated text, ensuring that the policies shaping the future are grounded in reality, not algorithmic imagination.[1]

How we got here

  1. 2023

    A study in Scientific Reports reveals that 55 percent of GPT-3.5 citations and 18 percent of GPT-4 citations are entirely fabricated.

  2. Late 2025

    The European cybersecurity agency ENISA publishes a threat landscape report containing 26 incorrect footnotes, later issuing a correction.

  3. Early 2026

    The South African government withdraws its draft national AI policy after media reports expose fabricated academic citations.

  4. April 2026

    Journalists discover 'oaicite' markers and fake studies in a flagship White House health assessment, prompting its retraction and revision.

  5. July 2026

    Policy experts at Stanford and the Carnegie Endowment formally classify AI-generated research fabrication as a national security vulnerability.

Viewpoints in depth

Policy Researchers' View

Academics and think-tank analysts view this as a systemic contamination of the information supply chain.

Researchers argue that because AI models generate text via statistical inference rather than factual retrieval, they are fundamentally unsuited for unverified policy drafting. This camp advocates for treating citation verification as mandatory infrastructure rather than an optional final check, warning that corrupted research inevitably leads to corrupted national security and regulatory decisions.

Government Agencies' View

Regulators and federal departments are caught between the mandate to modernize and the risk of institutional embarrassment.

While acknowledging the failures, government agencies emphasize that AI tools are necessary to process vast amounts of data quickly in a rapidly changing geopolitical landscape. Their focus is on implementing better human-in-the-loop review processes and automated verification tools before publication, rather than abandoning AI-assisted drafting entirely.

AI Developers' View

Model creators acknowledge the hallucination problem, particularly in niche domains where training data is sparse.

Developers argue that the solution lies in better integration of Model Context Protocols (MCPs) and Retrieval-Augmented Generation (RAG) to anchor models to live, verified databases. They caution users against relying on raw LLM outputs for factual citations, emphasizing that models are designed for language fluency, not as infallible search engines.

What we don't know

  • How many other local, state, and federal government documents currently in circulation contain undetected AI-generated citations.
  • Whether foreign adversaries are intentionally exploiting these vulnerabilities by poisoning the training data or search results that government AI tools rely upon.
  • How effectively upcoming automated citation-verification tools will be able to distinguish between highly sophisticated AI fabrications and obscure but legitimate research.

Key terms

AI Hallucination
A phenomenon where an artificial intelligence model generates false, misleading, or nonsensical information presented as fact.
oaicite marker
A hidden digital tag or fingerprint left in text generated by OpenAI's ChatGPT when it attempts to format a citation.
Statistical Inference
The process by which AI models predict the most likely sequence of words based on their training data, rather than retrieving specific factual records.
Model Context Protocol (MCP)
A technical standard that allows AI models to connect to external, live databases and tools to ground their responses in real-time data.

Frequently asked

Why do AI models invent fake citations?

AI models generate text by predicting the next most likely word based on statistical patterns, not by searching a database of facts. When they lack specific information, they often combine real author names and plausible journal titles to create convincing but entirely fabricated references.

Which government documents contained fake citations?

Fabricated citations were found in a White House health assessment, a threat landscape report by the European cybersecurity agency ENISA, and a draft national AI policy from the South African government.

How were the fake citations discovered?

Outside researchers and journalists identified the fabrications by checking the bibliographies, noting missing authors, incorrect formatting, and in the case of the White House report, the presence of 'oaicite' markers left behind by OpenAI's software.

Are newer AI models better at citing sources?

While newer models have improved, they still hallucinate. A 2025 analysis found that roughly 20 percent of citations generated by GPT-4o were fabricated, with the error rate increasing for highly specialized policy topics.

Sources

Source coverage

4 outlets

4 viewpoints surfaced

Policy Researchers & Academics 40%Government Agencies 30%AI Developers 15%Journalists & Watchdogs 15%
  1. [1]TechPolicy.PressPolicy Researchers & Academics

    When AI Hallucinates Policy: The National Security Threat of Fabricated Research

    Read on TechPolicy.Press
  2. [2]The Washington PostJournalists & Watchdogs

    White House Health Report Contained 'oaicite' Markers, Revealing AI Use

    Read on The Washington Post
  3. [3]OpenAIAI Developers

    Understanding Hallucinations and Citation Accuracy in GPT-4o

    Read on OpenAI
  4. [4]Carnegie Endowment for International PeacePolicy Researchers & Academics

    The Institutional Vulnerability of AI-Assisted Policy Research

    Read on Carnegie Endowment for International Peace
Stay informed

Every angle. Every day.

Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.