AI CopyrightLegal ExplainerJul 26, 2026, 6:18 AM· 5 min read

Publishers and Authors File Class-Action Lawsuit Against Google Over AI Copyright Infringement

Three major publishing houses and bestselling author Scott Turow have sued Google, alleging the tech giant illegally used millions of copyrighted books to train its Gemini AI models.

By Factlen Editorial Team

Publishers and Authors 45%AI Developers 35%Legal Analysts 20%
Publishers and Authors
Argue that using copyrighted books for AI training without permission or compensation is infringement that threatens the literary economy.
AI Developers
Maintain that training models on accessible data qualifies as fair use, similar to a human reading a book to learn.
Legal Analysts
Focus on the ambiguity of applying 20th-century copyright law to generative AI, noting the high financial stakes of the dispute.

What's not represented

  • · Independent Self-Published Authors
  • · Open-Source AI Researchers

Why this matters

The outcome of this lawsuit could redefine copyright law for the artificial intelligence era, determining whether tech companies must pay authors and publishers for the data used to train generative AI models.

Key points

  • Three major publishers and author Scott Turow filed a class-action lawsuit against Google over AI training data.
  • The lawsuit alleges Google used books provided for Google Play Books to train its Gemini AI without permission.
  • Internal Google documents allegedly warned that using publisher-provided books could result in massive fines.
  • The case tests whether generative AI training qualifies as "fair use" under U.S. copyright law.
$10Bs-$100Bs
Potential fines flagged in Google's internal warnings
3
Major publishers involved in the lawsuit
$1.5 billion
Anthropic's recent AI copyright settlement
2015
Year Google won its previous landmark book-scanning case

The intersection of artificial intelligence and copyright law has reached a critical new flashpoint. In July 2026, three of the world's largest publishers—Hachette Book Group, Cengage Learning, and Elsevier—joined forces with bestselling author Scott Turow to file a sweeping class-action lawsuit against Google.[1][2]

The complaint, filed in a New York federal court, accuses the technology giant of orchestrating "one of the most prolific infringements of copyrighted materials in history" to train its Gemini large language models. The plaintiffs are seeking class-action status to represent a vast coalition of authors and publishers whose works were allegedly ingested without permission.[1][3]

At the heart of the dispute is how Google allegedly sourced the massive volumes of text required to make Gemini conversational and intelligent. The publishers claim that Google repurposed millions of books and scholarly articles that had been originally provided to the company for strictly limited services, such as Google Books, Google Play Books, and Google Scholar.[4][6]

According to the plaintiffs, those initial agreements allowed Google to display searchable snippets or sell ebooks, but explicitly did not grant permission to ingest the texts into commercial AI training pipelines. By allegedly crossing that line, the lawsuit argues, Google transformed a mutually beneficial digital library partnership into an unauthorized data-mining operation.[3][4]

How large language models process and learn from written works.
How large language models process and learn from written works.

To understand the gravity of the publishers' claims, it is necessary to examine the mechanism of AI training. Large language models like Gemini do not "read" books in the human sense; they process them through a computational method called tokenization. The AI breaks down sentences into fragments, analyzing billions of linguistic patterns to predict which words should logically follow one another.[2][6]

For an AI to generate high-quality, coherent, and factually accurate text, it requires high-quality training data. While the open internet provides an endless supply of text, it is often poorly edited, repetitive, or factually dubious. Professionally edited books, memoirs, and peer-reviewed academic journals represent the "gold standard" of training data, teaching the model sophisticated grammar, narrative structure, and complex reasoning.[1][6]

The lawsuit alleges that Google knew the legal risks of tapping into this premium data. The complaint cites internal Google communications that allegedly flagged the use of "Publisher Provided" copyrighted books as "highly problematic." According to the filing, Google's own internal risk assessments warned that using Google Play Books for AI training could expose the company to "$10Bs-$100Bs in potential fines."[3][4]

The lawsuit alleges that Google knew the legal risks of tapping into this premium data.

Furthermore, the plaintiffs accuse Google of actively concealing its training sources by stripping Copyright Management Information (CMI)—the digital metadata that identifies authors and rights holders—from the ingested texts. This alleged removal of digital watermarks makes it significantly harder for authors to prove that their specific works were used to train the model.[7][8]

Google has historically defended its data-ingestion practices under the doctrine of "fair use," a provision in U.S. copyright law that allows limited use of protected works without permission for purposes like research or transformative creation. The company has successfully wielded this defense before, most notably in a decade-long legal battle over its original Google Books project.[7][8]

In 2015, the Second Circuit Court of Appeals ruled in Authors Guild v. Google that scanning millions of books to create a searchable database was a non-infringing fair use. The court determined that because Google only displayed short "snippets" of the books to users, the tool served a highly transformative purpose—acting as a digital card catalog—without providing a significant market substitute for the original works.[4][8]

However, legal experts note that generative AI presents a fundamentally different challenge than a search engine. While Google Books helped users discover and purchase the original texts, generative models like Gemini are designed to synthesize information and create entirely new content.[2][3]

The publishers argue that this crosses the line from discovery to substitution. The complaint explicitly warns that without proper guardrails, Gemini could generate a "100-page murder mystery set in a quiet seaside town filled with secrets" that directly competes with the original copyrighted novels it ingested. If an AI can summarize a textbook or mimic an author's distinct literary voice, it threatens the primary market for those works.[2][5]

This lawsuit is not occurring in a vacuum; it is part of a broader, industry-wide reckoning over AI and intellectual property. Over the past two years, authors and publishers have launched a barrage of legal challenges against major AI developers, including OpenAI, Meta, and Anthropic.[5]

Recent legal settlements indicate the rising financial stakes of AI copyright disputes.
Recent legal settlements indicate the rising financial stakes of AI copyright disputes.

The financial stakes of these disputes are staggering. In 2025, Anthropic agreed to a landmark $1.5 billion settlement in a class-action lawsuit brought by authors over the training of its Claude AI models. That settlement, which a federal judge recently finalized, established a precedent that AI companies may ultimately have to pay billions to resolve historical copyright claims.[1][5]

Unlike the Anthropic case, which centered heavily on allegations that the company downloaded books from known piracy sites like Library Genesis, the Google lawsuit focuses on the alleged misuse of legally acquired, publisher-provided files. This distinction tests a different legal boundary: whether a tech company can unilaterally expand the scope of a digital licensing agreement to include AI training.[2][6]

Courts must now decide whether 20th-century copyright laws apply to generative AI training.
Courts must now decide whether 20th-century copyright laws apply to generative AI training.

The publishing industry's ultimate goal is not necessarily to halt AI development, but to force a transition from unauthorized scraping to formal licensing. If the courts rule that AI training requires explicit permission and compensation for copyrighted works, the fundamental economics of artificial intelligence could be permanently altered.[2][4]

How we got here

  1. 2015

    Google wins a landmark fair-use case allowing it to scan books and display short snippets in search results.

  2. 2023-2024

    Authors begin filing a wave of copyright lawsuits against major AI developers, including OpenAI and Meta.

  3. 2025

    Anthropic agrees to a $1.5 billion settlement with authors over the training of its Claude AI models.

  4. July 2026

    Major publishers and author Scott Turow file a class-action lawsuit against Google over its Gemini AI.

Viewpoints in depth

Publishers and Authors

Argue that using copyrighted books for AI training without permission or compensation is infringement that threatens the literary economy.

The publishing industry views the unauthorized ingestion of books as an existential threat to the literary ecosystem. Authors argue that their professionally edited, meticulously crafted works are the "gold standard" that makes AI models conversational and intelligent. By using these works without permission, publishers claim tech companies are extracting the value of human creativity to build commercial products that could ultimately serve as market substitutes for the original books.

AI Developers

Maintain that training models on accessible data qualifies as fair use, similar to a human reading a book to learn.

Technology companies generally argue that training an AI model is fundamentally different from copying or distributing a book. They liken the machine learning process to a human student reading a library full of books to learn grammar, facts, and concepts. From this perspective, analyzing the statistical relationships between words is a highly transformative act that falls squarely under the "fair use" doctrine, and restricting it would stifle technological innovation.

Legal Analysts

Focus on the ambiguity of applying 20th-century copyright law to generative AI, noting the high financial stakes of the dispute.

Legal scholars emphasize that current copyright laws were not written with generative artificial intelligence in mind, leaving courts to interpret decades-old precedents in a radically new context. Analysts point out that while Google won its 2015 book-scanning case because its search engine did not create a market substitute for books, generative AI's ability to synthesize and create new text makes the fair-use defense much more complicated. The recent billion-dollar settlements indicate that the legal risk for AI companies is immense.

What we don't know

  • Whether the court will grant class-action status to the plaintiffs.
  • If Google will attempt to settle the lawsuit or take it to trial to establish a new legal precedent.
  • How a ruling against Google would impact the development and capabilities of future Gemini models.

Key terms

Large Language Model (LLM)
An AI system trained on massive amounts of text to understand and generate human language.
Fair Use
A legal doctrine that permits limited use of copyrighted material without permission for purposes like criticism, news reporting, or research.
Tokenization
The process of breaking down text into smaller units (tokens) that an AI model can process and learn from.
Copyright Management Information (CMI)
Metadata attached to a digital work that identifies the author, title, and copyright owner.

Frequently asked

Why are publishers suing Google now?

Publishers allege Google used books provided for limited services, like Google Play Books, to train its Gemini AI without permission or compensation.

Hasn't Google been sued over books before?

Yes. In 2015, Google won a landmark case over its Google Books project, with courts ruling that displaying short "snippets" of scanned books was fair use.

What do the authors want?

They are seeking class-action status, financial compensation for the unauthorized use of their work, and guardrails to prevent AI from generating substitute works.

How does this affect everyday readers?

The outcome could determine how AI models are trained in the future and whether authors receive royalties when their work contributes to AI generation.

Sources

Source coverage

8 outlets

3 viewpoints surfaced

Publishers and Authors 45%AI Developers 35%Legal Analysts 20%
  1. [1]Hachette Book GroupPublishers and Authors

    Publishers and Authors File Class Action Lawsuit Against Google for Willful Copyright Infringement

    Read on Hachette Book Group
  2. [2]The GuardianPublishers and Authors

    Group of major publishers accuses Google of 'prolific infringement' of copyrighted materials

    Read on The Guardian
  3. [3]EngadgetAI Developers

    Publishers and authors sue Google over Gemini AI training data

    Read on Engadget
  4. [4]TechRepublicAI Developers

    Google's AI copyright problems grow with new publisher lawsuit

    Read on TechRepublic
  5. [5]GizmodoLegal Analysts

    Anthropic's $1.5 Billion Copyright Settlement Finalized as AI Legal Battles Continue

    Read on Gizmodo
  6. [6]International Publishers AssociationPublishers and Authors

    Publishers and Authors File Class Action Lawsuit Against Google

    Read on International Publishers Association
  7. [7]JustiaLegal Analysts

    Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015)

    Read on Justia
  8. [8]US Copyright OfficeLegal Analysts

    Fair Use Index: Authors Guild v. Google, Inc.

    Read on US Copyright Office
Stay informed

Every angle. Every day.

Get culture stories with full source coverage and perspective breakdowns delivered to your inbox.