Skip to main content
AI Training DataAntitrust ExplainerAug 25, 2026, 11:20 AM· 4 min read· in ai

Civil Society Groups Urge FTC Antitrust Probe Over AI Firms Destroying Books for Training Data

A coalition of 18 advocacy organizations is petitioning the Federal Trade Commission to investigate major AI developers for systematically buying, digitizing, and destroying physical books to build private training datasets.

By Logan Price

Civil Society and Consumer Advocates 35%AI Developers and Incumbents 35%Legal and Antitrust Analysts 30%
Civil Society and Consumer Advocates
Argue that destroying physical books creates an anticompetitive moat and erases rare human knowledge.
AI Developers and Incumbents
View high-quality human text as an essential, legally acquirable resource for training competitive models.
Legal and Antitrust Analysts
Note that the shift from copyright to antitrust presents a novel regulatory challenge for the FTC.

Key terms

Destructive Scanning
The process of cutting the binding off a physical book to rapidly feed its loose pages through a high-speed optical scanner, resulting in the book's destruction.
Model Collapse
A degradation in AI performance that occurs when models are trained on synthetic, AI-generated data rather than high-quality human text.
Section 5 of the FTC Act
A U.S. federal law that prohibits unfair or deceptive acts or practices in commerce, including anticompetitive behavior.
Systemic Moat
A structural barrier to entry created by dominant companies that makes it prohibitively expensive or difficult for new competitors to enter a market.

Key points

  • A coalition of 18 civil society groups has petitioned the FTC to investigate AI companies for destructively scanning physical books.
  • The groups argue that hoarding and destroying rare, pre-2022 books creates an anticompetitive moat that starves startups of essential training data.
  • AI developers rely on older physical books to secure high-quality, human-written text that is guaranteed to be free of AI-generated content.
  • A 2025 federal court ruling previously established that destructively scanning legally purchased books constitutes fair use under copyright law.
  • The FTC petition attempts to shift the regulatory battleground from intellectual property disputes to antitrust enforcement.

In early 2024, an internal initiative at artificial intelligence startup Anthropic, codenamed "Project Panama," set out to acquire and digitize millions of physical books. To feed its Claude chatbot with high-quality human text, the company spent tens of millions of dollars buying books in bulk, slicing off their spines, running the loose pages through industrial scanners, and discarding the remains.

Now, that practice is facing federal scrutiny. A coalition of 18 civil society and consumer advocacy organizations has formally petitioned the U.S. Federal Trade Commission (FTC) to investigate major AI developers for systematically hoarding and destroying print books. The groups, which include the Demand Progress Education Fund and the Consumer Federation of America, argue that the mass destruction of physical media is not just an ethical concern, but a deliberate anticompetitive strategy.[1][2][5]

The core of the coalition's argument rests on Section 5 of the FTC Act, which prohibits unfair methods of competition. According to the petition, AI incumbents are buying up rare, out-of-print, and specialized titles, digitizing them for private training datasets, and permanently removing the physical copies from circulation. Critics argue this "hoard-and-destroy" tactic creates an insurmountable systemic moat, artificially raising costs for smaller startups that need the same high-quality source material to build rival systems.[1][3][5]

The sudden rush for physical books is driven by a phenomenon known as "model collapse." As the internet becomes increasingly saturated with AI-generated content—often referred to as "slop"—developers are desperate for pristine, human-written text. When an AI model trains on data generated by another AI, its outputs degrade over successive generations, losing nuance and structural integrity. To prevent this, frontier labs require massive volumes of text that predates the generative AI boom.[3][4]

Frontier AI models require pre-2022 human-written text to avoid the degrading effects of training on synthetic data.

Books published before 2022 are highly prized because they are guaranteed to be free of AI generation. They offer well-edited, structured knowledge that teaches models how to reason, maintain long-context coherence, and write effectively. Internal communications revealed in court filings show that developers consider books to be the highest-quality training material available, vastly superior to scraped web data or social media posts.[2]

Books published before 2022 are highly prized because they are guaranteed to be free of AI generation.

To acquire this data at scale, companies have turned to destructive scanning. The process is brutally efficient: hydraulic machines shear off the book's binding, allowing the loose pages to be fed rapidly through high-speed optical scanners. Once the text is converted into machine-readable data, the damaged paper is typically recycled or pulped.[3]

While some of the destroyed books are common mass-market paperbacks, advocates warn that the indiscriminate bulk-buying targets obscure local histories, old translations, and rare editions. In some cases, the petition alleges, these may be the last surviving copies of original works. The permanent removal of these texts from the public domain means that future researchers, historians, and competing AI developers can no longer access them.[1][2][4]

Once physical books are destructively scanned, the damaged remains are typically pulped or recycled.

The regulatory battle over this practice is shifting from intellectual property to antitrust because of recent legal precedents. In 2025, a federal judge ruled in a copyright lawsuit that Anthropic's use of legally purchased books to train its models constituted fair use. The court found that transferring the text from a physical medium to a digital one—without distributing the digital copies—did not violate copyright law.[1][3]

Crucially, the ruling established that destroying the original physical copy during the scanning process does not alter the fair use calculation. Because the AI developer legally purchased the physical object, they possess the right to destroy it, provided they do not sell or distribute the resulting digital reproduction. That ruling effectively closed one avenue of legal challenge, prompting civil society groups to pivot to competition law.[1][2][3]

By framing the destruction of books as an aggressive market tactic designed to starve competitors of essential inputs, the petitioners hope to force the FTC to intervene where copyright law has not. The groups argue that the practice violates Section 5 of the FTC Act, which bars unfair or deceptive acts in commerce. They contend that only the wealthiest incumbents can afford to buy and destroy millions of books, thereby engineering a future where they operate as the sole holders of humanity's written works.[1][2][5]

The process of converting physical books into machine-readable training data permanently removes the original copies from circulation.

If the FTC launches a formal probe, it could fundamentally reshape how the artificial intelligence industry sources its most valuable commodity. The agency would need to determine whether physical training data constitutes an essential market facility, and whether its systematic destruction actively harms market competition.[2][4]

For now, the practice remains a quiet but industrial-scale operation. Internal documents unsealed in the Anthropic lawsuit revealed that executives explicitly sought to keep the scanning efforts secret, acknowledging the controversial nature of dismantling physical books to build digital intelligence. As the FTC weighs the petition, the debate highlights a growing tension between the voracious data appetites of frontier AI models and the preservation of humanity's physical written record.[1][3]

Sources

Source coverage

5 outlets

3 viewpoints surfaced

Civil Society and Consumer Advocates 35%AI Developers and Incumbents 35%Legal and Antitrust Analysts 30%
  1. [1]CBS NewsCivil Society and Consumer Advocates

    AI companies accused of hoarding and destroying millions of books

    Read on CBS News
  2. [2]DataconomyLegal and Antitrust Analysts

    Civil society groups urge FTC to investigate AI firms for book destruction

    Read on Dataconomy
  3. [3]Ground NewsLegal and Antitrust Analysts

    AI Companies Are Burning Books, Advocates Complain to FTC

    Read on Ground News
  4. [4]OECD.AILegal and Antitrust Analysts

    AI Companies Accused of Destroying Books for Model Training, Prompting FTC Scrutiny

    Read on OECD.AI
  5. [5]AxiosCivil Society and Consumer Advocates

    Exclusive: FTC urged to investigate AI firms for destroying books

    Read on Axios

Comments

Stay informed

Every angle. Every day.

Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.