Skip to main content
ExplainerResearch IntegrityRetraction Watch· 7 min read· in Education

Why Retracted Scientific Papers Are Watermarked Rather Than Deleted

When a scientific paper is retracted for fraud or error, publishers do not erase it from their archives. Instead, the flawed research remains permanently indexed and watermarked to preserve the integrity of the historical scholarly record and prevent downstream citation cascades.

By Paige Carter

In short

  1. Scientific publishers watermark retracted papers rather than deleting them to ensure future researchers can verify why a foundational claim was invalidated.
  2. Retracted articles continue to accrue citations for years, largely because researchers read downloaded PDFs that lack the retroactive warning stamps.
  3. Plagiarized papers are retracted four times faster than those with fabricated data, yet they continue to be cited at double the rate post-retraction.

When a scientific paper is found to contain fabricated data or critical errors, it is not deleted from the internet. Instead, publishers leave the original document exactly where it is, permanently accessible but stamped with a bold "RETRACTED" watermark across every page.[6]

This counterintuitive practice exists to protect the historical scholarly record. If a flawed paper simply vanished, the hundreds of subsequent studies that cited it would point to a dead link, leaving future researchers unable to verify why a foundational claim was invalidated.[2]

The architecture of academic publishing relies on absolute permanence. Digital Object Identifiers (DOIs), the unique alphanumeric strings that link to research papers, are designed to resolve to a specific document forever.[5]

Registering a new DOI for content that already has one contravenes the membership terms of Crossref, the primary registration agency for scholarly literature. Once a DOI is minted, the publisher is obligated to maintain the landing page, even if the research itself is entirely discredited.[5]

The Committee on Publication Ethics (COPE) formalizes this approach in its global guidelines. COPE explicitly advises that retractions are issued to correct the scholarly record, not to punish authors, and that the original text must remain available to readers.[1]

The standard editorial workflow for investigating and retracting flawed scientific research.

Under these rules, a journal must issue a separate retraction statement that explains exactly why the paper was pulled. This notice is bi-directionally linked to the original article, ensuring that anyone who lands on the flawed research immediately sees the correction.[1]

The standard industry practice is codified in the editorial policies of thousands of journals, which explicitly define the scope of the corrective action.[1]

"Retraction is used when the published article's findings or conclusions can no longer be relied upon due to serious flaws," notes the editorial policy of Legal Research & Analysis. "Retractions apply to the Version of Record."[1]

The Scale of the Correction

The volume of retracted literature has grown substantially over the past two decades, driven by better detection tools and heightened scrutiny. In 2023 alone, more than 10,000 research papers were retracted across scientific publishing, setting a new annual record.[2]

The sheer scale of the scientific enterprise means that even a low baseline error rate produces a massive absolute number of flawed papers. With millions of articles entering the scholarly record annually, the infrastructure required to police, investigate, and flag misconduct has become a distinct industry of its own.[6]

A comprehensive bibliometric analysis of 46,087 retractions across ten major publishers revealed stark differences in how journals handle misconduct. Normalized retraction rates varied by two orders of magnitude, from 3.97 per 10,000 publications at Elsevier to 320.02 at Hindawi.[2]

Data concerns and outright fraud are the leading drivers of retractions in the medical literature.

The Retraction Watch database, which was acquired by Crossref in September 2023, now serves as the largest open-source repository of these corrections. The database catalogs tens of thousands of entries, providing a granular look at why science fails and how it recovers.[5]

Data concerns and outright fraud are the leading drivers of this corrective action. A 50-year analysis of medical retractions found that data issues accounted for 31.47 percent of cases, while fraud and peer-review manipulation each drove roughly 11 percent.[2]

Despite the rising numbers, retractions still represent a tiny fraction of the millions of papers published annually. However, their impact is disproportionately large because flawed findings can circulate in policy documents, clinical guidelines, and downstream research for years before being caught.[2]

The consequences of these delays are particularly severe in the biomedical sciences, where flawed research can directly influence clinical trial designs and patient care. A retracted paper that remains unchallenged for years can misdirect millions of dollars in grant funding toward dead-end hypotheses.[6]

The Post-Retraction Citation Problem

Leaving retracted papers online preserves transparency, but it also creates a persistent vulnerability in the scientific ecosystem. Empirical studies consistently show that retracted papers continue to accrue citations long after they have been officially invalidated by their publishers.[4]

A case study examining 987 articles retracted in 2014 found that the vast majority of subsequent citations treated the flawed research as legitimate. Researchers continued to use the invalidated findings to corroborate their own work, seemingly oblivious to the retraction notices.[4]

Illustration: Researchers often cite flawed papers because they rely on downloaded PDFs that lack retroactive warning stamps.

"Our results show that the vast majority of citations to retracted articles are positive despite of the clear retraction notice on the publisher's platform," the researchers noted. "Positive citations can be also seen to articles that were retracted due to ethical misconduct, data fabrication and false reports."[4]

The persistence of these citations highlights a breakdown in how scientific literature is consumed. Many researchers download PDFs to their personal reference managers, meaning they never see the watermark that gets added to the publisher's live website months or years later.[6]

The problem is compounded by the sheer volume of literature a modern researcher must synthesize. When authors cite dozens of papers in a single literature review, they rarely re-verify the publication status of every foundational text they read during their graduate studies.[6]

Furthermore, secondary databases and aggregators often fail to update their records promptly. If a scientist discovers a paper through a third-party search engine rather than the primary journal site, the crucial retraction flag may be entirely missing from the metadata.[5]

Divergence in Misconduct Types

The speed at which a paper is retracted, and its subsequent lifespan in the literature, depends heavily on the nature of the underlying misconduct. An analysis of 460 genetics articles retracted between 1970 and 2016 revealed a sharp divergence between plagiarism and fabrication.[3]

Plagiarized papers are caught much faster than fabricated ones, but they continue to accrue more citations after retraction.

Publishers are highly efficient at catching stolen text. The median time to retraction for a plagiarized genetics paper was just 1.3 years, largely because automated similarity-check software makes text overlap relatively easy to prove during editorial investigations.[3]

In contrast, uncovering fabricated or falsified data is a grueling, manual process. The median time to retraction for fabricated genetics research stretched to 4.8 years, allowing those papers to accumulate thousands of citations before the scientific record was finally corrected.[3]

Yet, the citation penalty for these offenses is inverted. While plagiarized papers are caught quickly, they continue to be cited heavily, with 42 percent of their total citations occurring post-retraction. Fabricated papers see a much lower post-retraction citation rate of 21.5 percent.[3]

This inverse relationship suggests a pragmatic, if ethically murky, calculus by the scientific community. Plagiarized papers often contain scientifically valid data that was simply stolen from another researcher, whereas fabricated data is fundamentally useless and cannot be replicated in subsequent experiments.[6]

Consequently, researchers who need the underlying data may continue to cite a plagiarized paper because the biological or chemical mechanism it describes actually works in the lab. In contrast, a paper built on fabricated data will inevitably fail to replicate, naturally depressing its long-term citation count.[6]

Automating the Scholarly Record

To combat the zombie-citation phenomenon, the publishing industry is increasingly turning to automated metadata alerts. Crossmark, a service operated by Crossref, embeds a standardized button on published articles that pings a central database to verify the document's current status.[5]

When a researcher clicks the Crossmark button on a PDF, the system checks whether the paper has been corrected, retracted, or subjected to an expression of concern. This bypasses the problem of static files sitting un-updated on a hard drive.[5]

Automated metadata alerts allow reference managers to ping a central database and warn researchers if a saved PDF has been retracted.

The integration of the Retraction Watch database into the Crossref REST API has further strengthened this safety net. Research institutions and reference-management software can now programmatically check a bibliography against the global retraction list in milliseconds.[5]

"For a research office, this means retraction status can be checked programmatically against a DOI rather than relying on a publisher notice being seen manually," notes the infrastructure documentation, fundamentally shifting the burden of verification from the human to the machine.[5]

Despite these technological advances, the core philosophy of academic publishing remains unchanged. The scholarly record is a permanent ledger of human inquiry, encompassing both our greatest breakthroughs and our most spectacular failures.[6]

Erasing a retracted paper would sanitize history, creating an illusion of flawless scientific progress while destroying the evidence needed to understand how errors propagate. By watermarking rather than deleting, science forces itself to confront its mistakes in plain sight.[6]

How we did this

Method
Comparison of post-retraction citation persistence and time-to-retraction across different categories of research misconduct (plagiarism versus fabrication/falsification).
What we found
An inverse relationship exists between detection speed and citation persistence: while publishers identify and retract plagiarized papers nearly four times faster than fabricated ones, the scientific community continues to cite plagiarized work at double the rate of fabricated work after retraction—likely because stolen data often remains scientifically valid, whereas fabricated data is fundamentally useless.
What we worked from
Limits of this analysis
The underlying data focuses on genetics articles, and raw citation counts do not distinguish between researchers citing a paper for its findings versus citing it as an example of misconduct.

Jargon, explained

Digital Object Identifier (DOI)
A unique, permanent alphanumeric string assigned to a published article that provides a persistent link to its location on the internet.
Version of Record
The final, peer-reviewed, and typeset version of a scientific article as it was formally published by the journal.
Crossmark
A digital button embedded in academic PDFs that connects to a central database to verify if the document has been updated, corrected, or retracted.
Paper Mill
An unethical commercial operation that fabricates scientific papers and sells authorship slots to researchers who need publications for career advancement.
Bibliometrics
The statistical analysis of written publications, often used to track citation patterns, publication output, and the impact of research.

Common questions

Can an author request to have their retracted paper completely deleted?

No. Once a paper is published and assigned a DOI, Crossref rules and COPE guidelines require the publisher to maintain the landing page and the watermarked PDF to preserve the historical record.

What is an 'Expression of Concern'?

It is a formal notice issued by a journal when serious doubts have been raised about a paper, but the investigation is either ongoing, inconclusive, or delayed. It serves as an interim warning to readers.

Do authors lose their degrees or jobs when a paper is retracted?

It depends entirely on the reason. Retractions for honest experimental errors carry little stigma, while retractions for deliberate data fabrication often trigger institutional investigations that can lead to dismissal.

How do reference managers handle retracted papers?

Historically, they did not flag them, which led to accidental citations. Recently, software like Zotero and EndNote have integrated with the Retraction Watch database to automatically warn users if a saved PDF is retracted.

Competing readings

Scientific Publishers

Publishers view the scholarly record as an immutable ledger that must reflect both successes and failures.

Publishers argue that deleting a flawed paper would create a 'dark archive' that hides the evidence of scientific failure. By keeping the Version of Record online with a retraction notice, they ensure that the hundreds of downstream papers that cited the flawed work do not suddenly point to a dead link, which would make it impossible to audit the spread of bad data.

Research Integrity Watchdogs

Integrity advocates argue that the current system is dangerously slow and allows fabricated data to circulate.

With investigations often taking years, fabricated data is allowed to influence clinical trials long before a watermark is applied. Watchdogs contend that publishers prioritize their own reputations over rapid correction, and that relying on a PDF watermark is inadequate when many researchers access literature through third-party aggregators that strip out the warning flags.

Working Researchers

Scientists rely on the permanence of the record but struggle to track retractions across static files.

For the scientists actually conducting experiments, the permanence of the record is a double-edged sword. A researcher might download a valid PDF in 2021, only for it to be retracted in 2023. Without automated tools pinging their personal reference managers, they remain entirely unaware that the foundational paper they are citing has been invalidated.

Scientific Publishers 40%Research Integrity Watchdogs 35%Working Researchers 25%
Scientific Publishers
Argue that maintaining a permanent, watermarked record is essential for transparency and preventing broken links in the web of citations.
Research Integrity Watchdogs
Argue that publishers move too slowly to investigate fraud and that watermarks are insufficient if secondary databases fail to display them.
Working Researchers
Value the preservation of the record but struggle to track retractions across thousands of static PDFs stored on personal hard drives.

Perspectives this story doesn't cover

  • Early-career researchers whose work was scooped by fabricated papers
  • University ethics committees tasked with investigating misconduct

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Scientific Publishers 40%Research Integrity Watchdogs 35%Working Researchers 25%
  1. [1]Legal Research & AnalysisScientific Publishers

    Retraction, Withdrawal, Correction, Removal, and Replacement Policies

    Read on Legal Research & Analysis →
  2. [2]arXiv

    Retractions and the Scholarly Record

    Read on arXiv →
  3. [3]Accountability in ResearchResearch Integrity Watchdogs

    For how long and with what relevance do genetics articles retracted due to research misconduct remain active in the scientific literature

    Read on Accountability in Research →
  4. [4]ScientometricsResearch Integrity Watchdogs

    Post retraction citations in context: a case study

    Read on Scientometrics →
  5. [5]CrossrefScientific Publishers

    Best practices for handling retractions and other post-publication updates

    Read on Crossref →
  6. [6]Factlen Editorial TeamWorking Researchers

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team →

Comments

Stay informed

Every angle. Every day.

Get Education stories with full source coverage and perspective breakdowns, free every day.