Anthropic Pays Record $1.5 Billion Copyright Settlement as Court Upholds Fair Use for AI Training
A federal judge has granted final approval to a landmark $1.5 billion settlement between Anthropic and 500,000 authors. The ruling establishes a critical legal precedent, determining that while training AI on books is fair use, acquiring those books through pirated shadow libraries is not.
By Factlen Editorial Team
- Authors & Rights Holders
- View the $1.5 billion payout as long-overdue accountability for rampant data theft, arguing that AI companies have built empires on uncompensated labor.
- AI Developers & Tech Industry
- Argue that training models is fundamentally transformative and protected by fair use, viewing the piracy penalty as a correctable sourcing error rather than a flaw in AI technology.
- Legal & Copyright Experts
- Emphasize the precarious nature of the split ruling, noting that while the distinction between training and acquisition solves this case, broader fair use questions remain unresolved.
What's not represented
- · Open-source AI developers
- · Shadow library operators
Why this matters
This ruling establishes the foundational legal framework for the generative AI industry. By separating the act of AI training (ruled legal) from the act of data acquisition (penalized for piracy), it forces tech companies to abandon unauthorized web scraping in favor of paid licensing agreements.
Key points
- U.S. District Judge Araceli Martínez-Olguín granted final approval to Anthropic's record $1.5 billion copyright settlement.
- The court ruled that training AI models on copyrighted text qualifies as fair use because it is highly transformative.
- However, the court penalized Anthropic for acquiring its training data through pirated shadow libraries like LibGen.
- The settlement will pay approximately $3,000 per eligible work to the rightsholders of 500,000 books.
The largest copyright payout in United States history has officially been finalized, fundamentally reshaping the legal framework governing artificial intelligence. On Monday, U.S. District Judge Araceli Martínez-Olguín granted final approval to a $1.5 billion settlement between AI developer Anthropic and a class of roughly 500,000 authors and publishers. The landmark resolution in Bartz v. Anthropic concludes a bitter two-year legal battle over the data used to train the company's Claude language models.[1][3]
The sheer scale of the settlement—amounting to roughly $3,000 per infringed work—is unprecedented in copyright litigation. Yet the most consequential outcome of the case is not the financial penalty, but the nuanced legal precedent it establishes for the entire generative AI industry. The court's rulings have effectively severed the computational act of AI training from the physical act of data acquisition, creating a complex new reality for developers and rightsholders alike.[3][6]
The lawsuit, originally filed in 2024 by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, accused Anthropic of "largescale theft" of copyrighted works. The plaintiffs argued that Anthropic built its multi-billion-dollar business by ingesting hundreds of thousands of books without permission or compensation. The case was closely watched as a bellwether for dozens of similar lawsuits pending against OpenAI, Meta, and Google.[1][3][5]
The critical turning point arrived in June 2025, when U.S. District Judge William Alsup issued a split summary judgment that stunned both sides. Addressing the core question of whether training an AI model on copyrighted text constitutes infringement, Alsup ruled decisively in favor of the tech industry. He declared that the use of books to train Claude was "transformative, spectacularly so," and thus protected under the fair use doctrine of U.S. copyright law.[5]

This finding of fair use for the actual training process was a massive victory for Anthropic and its peers. It affirmed the tech industry's central argument: that large language models do not simply copy and paste text, but rather analyze the statistical relationships between words to learn the underlying patterns of human language. If the ruling had stopped there, the authors would have lost their case entirely.[6]
However, Judge Alsup introduced a fatal caveat for Anthropic: the fair use protection only applies if the training data is acquired legally. The court found that Anthropic had populated its central training library by downloading millions of books from notorious "shadow libraries" like LibGen and PiLiMi—websites dedicated to pirating copyrighted material.[1][4]
"Anthropic had no entitlement to use pirated copies for its central library," the court concluded. "Creating a permanent, general-purpose library was not itself a fair use excusing Anthropic's piracy." This distinction meant that while the computational act of training the model was legal, the physical act of downloading and storing the pirated files to facilitate that training was a direct copyright violation.[5]
Facing a trial scheduled for December 2025 and the prospect of catastrophic statutory damages for willful infringement, Anthropic opted to settle. The company, which recently saw its valuation approach $1 trillion ahead of a rumored public listing, agreed to the $1.5 billion payout to eliminate the existential legal risk. The settlement received preliminary approval in September 2025 and final approval this week.[1][3][4]
Facing a trial scheduled for December 2025 and the prospect of catastrophic statutory damages for willful infringement, Anthropic opted to settle.
The mechanics of the payout are highly structured. The $1.5 billion fund will be distributed in four installments through September 2027. After deducting approximately $101.6 million in legal fees and administrative expenses, the remaining capital will be divided equally among the rightsholders of the eligible works.[1][4]

The settlement class encompasses approximately 500,000 distinct titles that were verified as having been downloaded from the shadow libraries and used in Anthropic's training runs. This broad inclusion means the payout will benefit a vast array of creators, from independent novelists to major academic publishers.[3][4]
Traditional publishing houses are already calculating their windfalls. Bloomsbury, the London-based publisher famous for the Harry Potter series, announced that 14,087 of its titles are included in the settlement agreement. The company expects to receive approximately $19 million after fees, which it plans to split with its authors starting in the second half of the fiscal year.[2]
For the authors who spearheaded the litigation, the final approval represents a hard-fought vindication. "Two years after we filed, the settlement for our class-action lawsuit got final approval," lead plaintiff Andrea Bartz wrote in a public statement following the ruling. "As I've been saying from the start, this is an important first step toward accountability for Big AI's breathtaking theft."[1]
Anthropic, meanwhile, is framing the conclusion of the case as a validation of its core technology. In a statement, Anthropic deputy general counsel Aparna Sridhar emphasized the fair use victory over the piracy penalty. "We reached this settlement in 2025, after the court's landmark ruling that training AI on books is fair use under copyright law—which remains the law today," Sridhar noted, adding that the company is pleased to bring the matter to a close.[2][3]
Legal analysts argue that Bartz v. Anthropic has fundamentally rewritten the playbook for AI data sourcing. The era of indiscriminately scraping the open web is effectively over. Data buyers and AI labs are now acutely aware that their legal exposure stems primarily from data provenance—how the information was acquired—rather than the algorithmic training process itself.[6]

This shift is already driving a massive surge in legitimate data licensing agreements. AI companies are rushing to sign multi-million-dollar contracts with news publishers, stock image libraries, and academic journals to secure clean, legally acquired training pipelines. The settlement proves that the cost of licensing data upfront is vastly cheaper than the penalty for pirating it later.[6]
Despite the clarity provided by the Anthropic settlement, the broader legal landscape remains fractured. While federal judges in California have leaned toward finding AI training to be fair use, courts in other jurisdictions have signaled skepticism. In Thomson Reuters v. Ross Intelligence, a federal court rejected a fair use defense for an AI legal search tool, noting that the AI's output competed directly with the original copyrighted material.[6]

The ultimate test of the fair use doctrine will likely come from the ongoing, high-stakes litigation in the Southern District of New York, where The New York Times and a coalition of other publishers are suing OpenAI and Microsoft. That case involves not just the ingestion of data, but allegations that the AI models are capable of reproducing near-verbatim copies of paywalled articles, directly threatening the publishers' core business.[1]
Until those cases reach appellate courts or the Supreme Court, the Anthropic settlement stands as the definitive ruling of the generative AI era. It establishes a precarious compromise: AI models can legally learn from human knowledge, but the companies building them can no longer pretend that the digital libraries they rely on are free for the taking.
How we got here
2024
Authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson file a class-action lawsuit against Anthropic.
June 2025
Judge William Alsup rules that training AI is fair use, but acquiring books from pirated libraries is not.
September 2025
Anthropic agrees to a $1.5 billion settlement to avoid a trial over the piracy claims.
July 2026
Judge Araceli Martínez-Olguín grants final approval to the record-breaking settlement.
Viewpoints in depth
AI Developers' View
The tech industry views the fair use ruling on training as a massive, existential victory.
AI developers argue that if learning from data was outlawed, the United States would lose its competitive edge in artificial intelligence. They view the piracy penalty as an avoidable sourcing error rather than a fundamental flaw in AI technology. By establishing that the computational analysis of text is transformative, the ruling provides the legal cover necessary for the continued development of frontier models, provided the data pipelines are clean.
Authors and Publishers' View
Creators see the $1.5 billion payout as long-overdue accountability for rampant data theft.
Rightsholders argue that AI companies have built trillion-dollar empires by strip-mining human creativity without offering compensation. They view this settlement as proof that copyright law still protects their labor in the digital age. While disappointed that the act of training itself was deemed fair use, publishers are leveraging the piracy penalty to force AI labs into lucrative, multi-million-dollar licensing agreements.
Legal Analysts' View
Copyright experts emphasize the precarious nature of the split ruling.
Legal scholars note that while the distinction between training and acquisition solves the immediate Anthropic case, it leaves unresolved the harder questions of market substitution. They point out that if legally acquired AI outputs begin to directly compete with and replace the original authors' market—such as AI generating new chapters in the style of a specific novelist—the fair use defense could still collapse in future appellate rulings.
What we don't know
- How appellate courts will rule if the split decision between data training and data acquisition is challenged in future cases.
- Whether the ongoing lawsuit between The New York Times and OpenAI will reach a similar conclusion regarding fair use.
- How the settlement will impact the training costs and data acquisition strategies of open-source AI developers who lack billion-dollar budgets.
Key terms
- Fair Use
- A legal doctrine that allows the unlicensed use of copyright-protected works in certain circumstances, such as when the use is highly transformative.
- Shadow Library
- An online database of content that is normally obscured or otherwise not readily accessible, often containing pirated copies of books and academic papers.
- Transformative Use
- A key factor in fair use analysis that asks whether the new work adds new expression, meaning, or message to the original material.
- Summary Judgment
- A legal decision made by a court without a full trial, typically when there is no dispute about the key facts of the case.
Frequently asked
Does this mean AI companies can no longer train on copyrighted books?
No. The court actually ruled that training AI on copyrighted text is 'fair use.' The penalty was specifically for how the books were acquired—downloaded from pirated shadow libraries.
How much will authors receive from the settlement?
Eligible authors and publishers will receive approximately $3,000 per work, divided among roughly 500,000 verified titles.
Does this ruling apply to all AI companies?
While it sets a massive precedent, it is a district court ruling and a settlement, not a Supreme Court mandate. Other courts are still deliberating similar cases, such as The New York Times' lawsuit against OpenAI.
Sources
[1]Los Angeles TimesAuthors & Rights Holders
Anthropic to pay $1.5 billion in record AI copyright settlement
Read on Los Angeles Times →[2]The GuardianAuthors & Rights Holders
Bloomsbury among beneficiaries of $1.5bn Anthropic copyright settlement
Read on The Guardian →[3]Silicon RepublicAI Developers & Tech Industry
US judge approves Anthropic's $1.5bn copyright settlement
Read on Silicon Republic →[4]The Authors GuildAuthors & Rights Holders
Anthropic Copyright Lawsuit Settlement Details
Read on The Authors Guild →[5]BakerHostetlerLegal & Copyright Experts
AI Copyright Litigation Update: Bartz v. Anthropic
Read on BakerHostetler →[6]AI Copyright LegalLegal & Copyright Experts
AI Training and Fair Use: What the Law Actually Says in 2026
Read on AI Copyright Legal →
More in ai
See all 5 stories →AI Compute
AMD Invests $5 Billion in Anthropic Under New AI Infrastructure Deal, Directly Challenging Nvidia
8 sources
Agentic AI
OpenAI Autonomous Agent Escapes Testing Sandbox and Hacks Hugging Face Infrastructure
2 sources
AI Prompting
New 'Seed-of-Thought' Prompting Technique Solves AI's Core Randomness Problem
6 sources
Every angle. Every day.
Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.










