Skip to main content
AI Copyright LawsuitsLegal Discovery· 4 min read· in Entertainment

OpenAI and Microsoft Executives Knew Book Piracy Was Illegal, Unsealed Court Filings Reveal

Newly unsealed internal communications show that OpenAI and Microsoft leadership acknowledged using pirated books to train AI models and anticipated the economic disruption to human authors.

By Dmitry Volkov

Authors and Publishers 40%Tech Industry Watchdogs 30%Legal and Media Analysts 30%
Authors and Publishers
Creators argue that the tech companies knowingly built commercial empires on stolen labor.
Tech Industry Watchdogs
Observers highlighting the internal ethical conflicts and corporate practices within AI development.
Legal and Media Analysts
Commentators focusing on the fair use defense and the broader legal implications of the copyright lawsuits.

Perspectives this story doesn't cover

  • Independent AI researchers
  • Open-source model developers

Why this matters

The unsealed communications provide the first direct evidence that AI developers internally acknowledged their training methods relied on pirated material and would economically damage human creators, potentially shifting the balance in ongoing copyright litigation.

Key points

  • Newly unsealed court filings reveal OpenAI and Microsoft executives knew their AI models were trained on pirated books.
  • Internal messages show engineers described the Library Genesis data source as 'sketchy AF' in 2019.
  • A Microsoft director described the mass scraping of copyrighted material as the 'largest theft of labor in human history.'
  • OpenAI leadership internally acknowledged that their generative models would cause 'acceptable economic disruption' to human authors.

The discovery phase of a federal lawsuit is where the polished corporate narrative goes to die, replaced by the blunt reality of internal Slack messages. Newly unsealed filings in the copyright battle against OpenAI and Microsoft have just dragged those private conversations into the public record, revealing exactly what executives were saying behind closed doors. Documents released in Manhattan federal court show that leadership at both companies internally acknowledged their artificial intelligence models were trained on pirated books, fully aware that the technology would directly threaten the livelihoods of human authors.[1][4]

The disclosures stem from consolidated copyright litigation brought by the Authors Guild, alongside prominent writers including John Grisham and George R.R. Martin, as well as news publishers like The New York Times. The plaintiffs are seeking summary judgment, arguing that the tech giants committed willful copyright infringement by ingesting millions of protected works to build systems like ChatGPT.[1][5]

Central to the authors' case is the revelation that early versions of OpenAI's language models were trained on data from Library Genesis, or LibGen—a widely known shadow library that hosts pirated books. It turns out the engineers building the future of text generation knew exactly where their material was coming from. According to the unsealed briefs, OpenAI engineers Tom Brown and Ben Mann internally described the data source as "sketchy AF" back in 2019.[4][5]

The internal communications show a deliberate effort to obscure the origins of this training data. When Dario Amodei, then a senior researcher at OpenAI and now the chief executive of Anthropic, asked if it was problematic to name the training corpora "Books1" and "Books2" without disclosing their source, Mann replied that the description was "deliberately vague since it's LibGen."[5]

Internal messages show engineers acknowledged using pirated datasets to train early language models.

The filings also highlight executives' awareness of the economic impact their products would have on the publishing industry. In May 2020, OpenAI policy director Jack Clark cautioned that improved language models would increasingly displace genre fiction writers, predicting the company would proceed despite pushback from artists.[4]

The filings also highlight executives' awareness of the economic impact their products would have on the publishing industry.

Two years later, the internal ambition had only grown. OpenAI research lead Tarun Gogineni outlined goals for autocompleting stalled literary series—a wry nod to famously delayed fantasy epics. When discussing authors' complaints that their work was being stolen to create AI-generated competition, Gogineni reportedly didn't mince words, dismissing the concerns as "acceptable economic disruption."[4]

Microsoft, which heavily invested in OpenAI, was also aware of the practices. The documents indicate that Microsoft executives, including former Chief Technology Officer Kevin Scott and co-founder Bill Gates, were informed about the use of LibGen as early as April 2019 during a walkthrough of an early GPT-3 model.[4]

Microsoft's own leadership expressed alarm at the sheer scale of the data ingestion. Brent Hecht, Microsoft's director of applied science, didn't hold back in his internal assessments. He described the mass scraping of copyrighted material as an "astonishing theft of unprecedented proportions" and suggested it could be the "largest theft of labor in human history."[1]

The Authors Guild argues that unchecked AI training poses an existential threat to the publishing industry.

OpenAI eventually initiated a cleanup effort in the summer of 2022. Then-vice president of research Bob McGrew directed the deletion of the LibGen-derived datasets, noting that removing the material from their systems would be "very valuable for legal reasons" as the company's public profile grew.[5]

The tech companies maintain that their training methods are protected under the legal doctrine of fair use, arguing that the models analyze information to generate transformative new responses rather than simply reproducing the original texts. They also note that the employees involved in creating the LibGen datasets are no longer with the company and that current models do not rely on that specific data.[1][5]

The unsealed documents do not resolve the underlying legal question of whether training artificial intelligence on copyrighted material constitutes infringement. However, the internal admissions provide the plaintiffs with direct evidence of corporate intent, shifting the weight of the upcoming trial as both sides await a ruling on their competing motions for summary judgment.[1][2]

How we got here

  1. April 2019

    Microsoft executives are informed that early GPT models are being trained on data from the piracy site Library Genesis.

  2. May 2020

    OpenAI policy director Jack Clark internally predicts that language models will increasingly displace genre fiction writers.

  3. Summer 2022

    OpenAI deletes the LibGen-derived datasets, with executives citing the move as valuable for legal reasons.

  4. Late 2023

    The Authors Guild and The New York Times file separate copyright infringement lawsuits against OpenAI and Microsoft.

  5. September 2026

    Federal courts unseal internal communications as part of the plaintiffs' motion for summary judgment.

Viewpoints in depth

The Authors and Publishers

Creators argue that the tech companies knowingly built commercial empires on stolen labor.

The Authors Guild and news publishers contend that the unsealed communications prove willful infringement. They argue that OpenAI and Microsoft intentionally bypassed paywalls and utilized known piracy sites like LibGen to avoid paying for training data, fully aware that their generative models would eventually compete with and displace the human creators who produced the original works.

The AI Developers

Tech companies maintain that their data ingestion practices are legally protected and transformative.

OpenAI and Microsoft argue that training artificial intelligence on publicly available or acquired data falls under the fair use doctrine. They assert that their models analyze texts to learn patterns and generate novel responses, rather than reproducing copyrighted material. OpenAI has also emphasized that the specific LibGen datasets in question were deleted in 2022 and are not used in current iterations of ChatGPT.

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Authors and Publishers 40%Tech Industry Watchdogs 30%Legal and Media Analysts 30%
  1. [1]KarmactiveLegal and Media Analysts

    OpenAI Microsoft Copyright Lawsuit Unsealed Filings Largest Theft Labor

    Read on Karmactive →
  2. [2]Complete AI TrainingTech Industry Watchdogs

    Unsealed filings show OpenAI and Microsoft executives knew their book piracy was illegal and would displace authors

    Read on Complete AI Training →
  3. [3]MediaPostLegal and Media Analysts

    Unsealed Court Filing Shows OpenAI, Microsoft Supplant Publishers' Sites 09/18/2026

    Read on MediaPost →
  4. [4]The Authors GuildAuthors and Publishers

    Unsealed Briefs in Authors' Case v. Microsoft/OpenAI: Top Execs Knew Their Mass Book Piracy Was Illegal And Would Put Authors Out of Work

    Read on The Authors Guild →
  5. [5]The BooksellerAuthors and Publishers

    OpenAI staff discussed using books obtained from Library Genesis

    Read on The Bookseller →

Comments

Stay informed

Every angle. Every day.

Get Entertainment stories with full source coverage and perspective breakdowns delivered to your inbox.