Skip to main content
AI Copyright LawIndustry LawsuitAug 10, 2026, 4:56 AM· 5 min read· #3 of 4 in entertainment

Publishers and Authors File Class-Action Lawsuit Against Google Over Gemini AI Training

A coalition of major book publishers and authors has sued Google, alleging the tech giant illegally copied millions of copyrighted works to train its Gemini artificial intelligence models.

By Chen Wang

Publishing Industry 40%AI Developers 30%Legal Observers 30%
Publishing Industry
Argues that tech companies are illegally exploiting copyrighted works to build commercial AI products that threaten authors' livelihoods.
AI Developers
Maintains that training artificial intelligence on digitized text is a transformative fair use protected by copyright law.
Legal Observers
Focuses on the shifting legal landscape, noting that this New York filing tests the boundaries of fair use following recent settlements.

The tension between Silicon Valley's insatiable appetite for data and the publishing industry's right to control its intellectual property has finally spilled over into a New York federal court. For years, tech giants have scraped the internet to train the artificial intelligence models that now write code, draft emails, and generate novels on command. But a coalition of major book publishers and authors is drawing a hard line, arguing that the era of "move fast and ingest everything" must end. On July 10, 2026, Hachette Book Group, Cengage Learning, Elsevier, and bestselling author Scott Turow filed a class-action lawsuit against Google, accusing the company of executing "one of the most prolific infringements of copyrighted materials in history" to build its Gemini AI models.[1][4]

The lawsuit strikes at the core of how modern generative AI is developed. According to the complaint, Google did not simply scrape the open web; it allegedly repurposed millions of books and journal articles that publishers had provided for strictly limited services, such as Google Books and Google Play Books. Those agreements were designed to allow searchable snippets or digital retail distribution, not to serve as the foundational training data for a commercial AI product. By feeding these curated, professionally edited texts into Gemini, the plaintiffs argue, Google bypassed the very contracts it signed to gain access to the material in the first place.[2][3][5][7]

What makes this particular filing explosive is the inclusion of internal Google communications. The publishers' complaint cites documents suggesting that Google engineers were acutely aware of the legal risks involved in using publisher-provided texts. According to the filing, internal memos flagged the practice as "highly problematic," warning that secretly training on Google Play Books could expose the company to anywhere from $10 billion to $100 billion in potential fines. The documents allegedly show Google acknowledging that publishers were highly sensitive about their data and would likely view the unauthorized ingestion as blatant copyright infringement.[1][6][7]

The lawsuit alleges Google repurposed books provided for digital retail to train its Gemini AI models.
The lawsuit alleges Google repurposed books provided for digital retail to train its Gemini AI models.

Beyond the breach of contract claims, the lawsuit accuses Google of actively covering its tracks. The plaintiffs allege that Google violated the Digital Millennium Copyright Act by stripping or altering copyright management information from the digital files before feeding them into the Gemini training algorithms. This removal, the publishers argue, was a deliberate attempt to obscure the origins of the training data and hide the fact that copyrighted works were being used to teach the AI how to mimic human writing styles and synthesize complex information.[2][3][4]

Beyond the breach of contract claims, the lawsuit accuses Google of actively covering its tracks.

The stakes for the publishing industry are existential. Generative AI models like Gemini are not just search engines; they are capable of producing long-form text that directly competes with the authors whose work trained them. The lawsuit notes that Gemini can generate a 100-page murder mystery set in a quiet seaside town in a matter of minutes, effectively substituting for the original copyrighted novels it ingested. For authors like Scott Turow, whose legal thrillers rely on intricate plotting and curated facts, the prospect of an AI instantly generating knockoff narratives represents a direct threat to their livelihoods.[1][4][5]

Google has historically defended its data collection practices under the doctrine of "fair use," arguing that training AI on copyrighted text is a transformative process that creates entirely new tools rather than derivative works. The company previously won a landmark copyright case regarding the digitization of books for its search engine, where a federal appellate court ruled that displaying snippets of text was legally permissible. However, the publishers contend that ingesting complete works to train a commercial generative AI system that competes with the original authors falls far outside the boundaries of that precedent.[4][6]

Publishers argue that generative AI models directly compete with the authors whose work was used to train them.
Publishers argue that generative AI models directly compete with the authors whose work was used to train them.

The choice of venue is also a strategic calculation. While recent California court rulings have broadly favored AI companies by treating training data ingestion as fair use, this lawsuit was filed in the U.S. District Court for the Southern District of New York. Legal observers note that the New York courts may offer a fresh judicial interpretation of how copyright law applies to generative AI, potentially setting up a circuit split that could eventually force the Supreme Court to weigh in on the issue.[2][4]

This lawsuit is part of a broader, industry-wide reckoning over AI training data. It follows a wave of similar litigation against companies like OpenAI, Meta, and Anthropic. Just weeks before this filing, a federal judge finalized a landmark $1.5 billion class-action settlement between Anthropic and hundreds of thousands of writers over the unlicensed use of digitized books to train its Claude chatbot. That settlement has emboldened the publishing industry, signaling that courts are increasingly willing to hold AI developers financially accountable for the data they consume.[1][2][5]

As the case moves forward, it will test the legal and economic foundations of the generative AI boom. If the publishers succeed in securing an injunction or massive statutory damages, it could force tech giants to fundamentally alter how they source their training data, shifting the industry toward a licensing model where creators are compensated for their contributions. Until then, the lawsuit stands as a stark declaration from the publishing world: the raw material of the AI revolution is not free for the taking.

The stakes

This lawsuit tests the legal boundaries of how artificial intelligence is developed, challenging the tech industry's reliance on free, scraped data. If the publishers prevail, it could force AI companies into a costly licensing model, fundamentally altering the economics of generative AI and setting a new precedent for how creators are compensated.

The essentials

  • Hachette, Cengage, Elsevier, and author Scott Turow have filed a class-action lawsuit against Google in New York.
  • The plaintiffs allege Google used millions of copyrighted books to train its Gemini AI models without authorization or compensation.
  • The lawsuit claims Google bypassed limited-use agreements for services like Google Play Books to access the training data.
  • Internal Google documents cited in the complaint allegedly warned of massive potential fines for using publisher-provided texts.
  • The publishers argue that generative AI poses an existential threat by producing content that directly competes with human authors.

Timeline

  1. 2015

    Google wins a landmark copyright case allowing the digitization of books for searchable snippets under fair use.

  2. 2023-2024

    A wave of copyright lawsuits is filed against AI developers, including OpenAI, Meta, and Anthropic, over training data.

  3. June 2026

    Anthropic agrees to a $1.5 billion class-action settlement with authors over the unlicensed use of digitized books.

  4. July 10, 2026

    Major publishers and author Scott Turow file a class-action lawsuit against Google in New York federal court.

Perspectives explored

The Publishers' Argument

Publishers contend that Google violated limited-use contracts to build a competing commercial product.

The core of the plaintiffs' argument rests on the assertion that Google abused its existing relationships with the publishing industry. By taking books provided specifically for retail distribution on Google Play Books and repurposing them as training data for Gemini, the publishers argue Google committed willful copyright infringement. They emphasize that generative AI models are not mere search tools but commercial products capable of generating substitute works that directly threaten authors' livelihoods, making the unauthorized ingestion a severe market threat.

The Fair Use Defense

Tech companies generally argue that training AI on existing text is a transformative process protected by copyright law.

While Google has yet to formally respond to this specific complaint in court, the technology industry's standard defense relies heavily on the doctrine of fair use. AI developers argue that machine learning algorithms do not copy or reproduce the books they ingest; rather, they analyze the text to learn patterns, grammar, and facts—a highly transformative use. They point to previous legal victories, such as Google's 2015 win regarding the digitization of books for search snippets, as precedent that using copyrighted works to build entirely new, innovative tools serves the public interest and does not violate copyright law.

Sources

Source coverage

7 outlets

3 viewpoints surfaced

Publishing Industry 40%AI Developers 30%Legal Observers 30%
  1. [1]The GuardianPublishing Industry

    Group of major publishers accuses the tech giant of 'one of the most prolific infringements of copyrighted materials in history'

    Read on The Guardian
  2. [2]The Next WebAI Developers

    Publishers sue Google over Gemini AI training

    Read on The Next Web
  3. [3]EngadgetAI Developers

    A trio of publishers and one author are seeking a class action lawsuit against Google

    Read on Engadget
  4. [4]BetaNewsLegal Observers

    Publishers and authors file class-action lawsuit against Google over Gemini AI

    Read on BetaNews
  5. [5]HypebeastLegal Observers

    Hachette Book Group, other publishers and author Scott Turow have sued Google

    Read on Hypebeast
  6. [6]Hachette Book GroupPublishing Industry

    Publishers and Authors File Class Action Lawsuit Against Google for Willful Copyright Infringement to Develop Gemini AI Models

    Read on Hachette Book Group
  7. [7]Association of American PublishersPublishing Industry

    Publishers and Authors File Class Action Lawsuit Against Google

    Read on Association of American Publishers

Comments

Stay informed

Every angle. Every day.

Get entertainment stories with full source coverage and perspective breakdowns delivered to your inbox.