Skip to main content
AI CopyrightLegal Battle· 5 min read· in Artificial Intelligence

Major Publishers and The New York Times Intensify Legal Battle Over AI Training Data

Nearly 400 local newspapers have filed a sweeping copyright lawsuit against OpenAI and Microsoft, while The New York Times amended its ongoing case to directly target Microsoft's supercomputing infrastructure.

By Harper Lane

News Publishers 40%AI Developers 40%Legal Analysts 20%
News Publishers
Argue that unauthorized AI training is mass digital theft that threatens the economic survival of journalism.
AI Developers
Maintain that training models on public data is protected fair use and critical for technological advancement.
Legal Analysts
View the lawsuits as a necessary stress test to adapt 20th-century copyright law to the generative AI era.

Perspectives this story doesn't cover

  • Open-Source AI Developers
  • Independent Content Creators and Artists

Key terms

Large Language Model (LLM)
An artificial intelligence system trained on vast amounts of text data to understand and generate human-like language.
Fair Use
A legal doctrine in U.S. copyright law that permits limited use of copyrighted material without permission for purposes such as criticism, news reporting, teaching, and research.
Contributory Infringement
A legal concept where a party can be held liable for copyright infringement if they intentionally induce or encourage another party to commit the illegal act.
Digital Millennium Copyright Act (DMCA)
A 1998 U.S. law that criminalizes the circumvention of digital access controls and the removal of copyright management information.

Key points

  • A coalition of nearly 400 local newspapers filed a copyright infringement lawsuit against OpenAI and Microsoft.
  • The publishers allege the tech companies scraped paywalled articles and stripped out copyright management data.
  • The New York Times amended its 2023 lawsuit to directly target Microsoft's role in building AI infrastructure.
  • The Times claims Microsoft built a custom supercomputer specifically to ingest copyrighted works at scale.
  • AI companies defend their data scraping practices as transformative fair use essential for technological progress.
  • The rulings could force AI developers to license training data or fundamentally alter how models are built.

The legal foundation of the generative artificial intelligence industry is facing its most coordinated and aggressive challenge to date. In a span of just 24 hours, a coalition of nearly 400 local newspaper publishers launched a sweeping federal lawsuit against OpenAI and Microsoft, while The New York Times simultaneously escalated its own ongoing copyright infringement case against the two tech giants. These parallel legal maneuvers, filed in the U.S. District Court for the Southern District of New York, signal a strategic shift in how the media industry is battling unauthorized AI training, moving from broad complaints about model outputs to highly technical attacks on the underlying infrastructure.[3]

The new class-action lawsuit is led by the Local News Copyright Alliance, a group representing independent and locally owned publishing companies across 33 states, including properties like the Tahoe Daily Tribune and the Arkansas Democrat-Gazette. The coalition alleges that OpenAI and Microsoft systematically bypassed paywalls and digital access restrictions to copy hundreds of thousands of articles. According to the complaint, the tech companies used this vast repository of local journalism to train models like ChatGPT and Microsoft Copilot without ever seeking permission or offering compensation to the original creators.[1]

The local publishers argue that this unauthorized scraping constitutes a 'death knell' for an already fragile local news industry. By feeding expensive, labor-intensive local reporting into chatbots that summarize the news directly for users, the publishers claim the tech companies are actively cannibalizing the subscription and advertising revenues required to fund original journalism. The lawsuit emphasizes that local news remains one of the most trusted sources of information in America, and that its unchecked exploitation by trillion-dollar tech companies threatens the civic fabric of communities nationwide.[1]

Crucially, the local publishers' lawsuit introduces a potent new legal argument under the Digital Millennium Copyright Act (DMCA). The plaintiffs allege that when OpenAI utilized automated text extraction tools—internally referred to as 'Dragnet' and 'Newspaper'—it deliberately stripped out vital copyright management information. This allegedly included author bylines, copyright notices, and terms of use. If proven in court, this DMCA violation carries separate and significant statutory penalties that go beyond standard copyright infringement, potentially exposing the tech companies to massive financial liabilities regardless of their fair use defenses.

The legal challenges highlight both the volume of publishers involved and the massive computing infrastructure required for AI training.

Simultaneously, The New York Times filed a heavily redacted motion to amend its December 2023 lawsuit, pivoting its legal crosshairs directly onto Microsoft. Previously viewed primarily as OpenAI's passive cloud provider and financial backer, Microsoft is now accused by the Times of actively inducing and facilitating mass copyright infringement. The amended complaint attempts to shift Microsoft's role from a neutral landlord providing generic server space to an active participant that knowingly engineered the machinery required to copy the world's news at an unprecedented scale.[3]

Simultaneously, The New York Times filed a heavily redacted motion to amend its December 2023 lawsuit, pivoting its legal crosshairs directly onto Microsoft.

The Times adapted its legal strategy following a recent Supreme Court ruling involving Cox Communications, which established a significantly tighter standard for 'contributory infringement.' Under the new legal precedent, plaintiffs can no longer simply prove that a party knew copyright infringement was happening on its servers. Instead, they must demonstrate that the party intentionally acted to induce the illegal conduct. Recognizing this shifted landscape, the Times moved quickly to realign its complaint to meet the higher burden of proof required to hold Microsoft accountable.[2]

To meet this higher legal bar, the Times alleges that Microsoft did not merely rent out off-the-shelf cloud servers to OpenAI. Instead, the amended complaint claims Microsoft built a bespoke, 'unusually complex' supercomputing system specifically designed to ingest copyrighted material. According to the filings, this custom infrastructure reportedly features over 285,000 CPU cores and 10,000 GPUs, representing a massive capital investment explicitly tailored to consume the entire internet and train the most capable large language models in history.[2]

The Times argues this infrastructure was curated to disproportionately feature high-quality journalism, ensuring that the resulting AI models could confidently mimic professional writing styles. By framing Microsoft as the architect of this ingestion engine, the lawsuit attempts to pierce the liability shield traditionally enjoyed by cloud infrastructure providers. If a court treats the provision of AI supercomputing infrastructure as active participation in scraping, every major cloud company hosting AI models could face profound new legal risks.[2]

The New York Times alleges Microsoft built a bespoke supercomputing system specifically designed to ingest copyrighted material at scale.

OpenAI and Microsoft have consistently defended their data collection practices under the doctrine of 'fair use,' a cornerstone of U.S. copyright law that permits limited use of copyrighted material for transformative purposes. The companies maintain that training AI models on publicly available data is fundamentally transformative because the models learn statistical language patterns and facts, rather than simply storing and reproducing copyrighted expression verbatim. They argue this process is akin to a human student reading a library of books to learn how to write.[1][2]

OpenAI representatives have previously stated that ChatGPT is not a substitute for a news subscription and that restricting the use of public internet data would severely hinder technological innovation and national competitiveness. Following the recent filings, Microsoft characterized the Times' amended complaint as a 'last-ditch effort' to save its legal claims from unfavorable precedents set in other recent rulings, maintaining that their partnership with OpenAI operates entirely within the bounds of established copyright law.[2]

The outcome of these consolidated legal battles will likely dictate the future economics and technical architecture of artificial intelligence. If the courts ultimately side with the publishers and reject the fair use defense, AI companies could be forced to negotiate massive licensing deals with content creators across the globe. In the most extreme scenario, courts could order tech companies to disgorge previous profits or even wipe their existing models and rebuild them entirely from licensed data, fundamentally resetting the AI arms race.[2]

Why this matters

These parallel lawsuits will likely determine the economic foundation of the generative AI industry. If courts rule that AI developers must license all training data, it could force a massive transfer of wealth from tech companies to publishers and fundamentally alter how future AI models are built.

What we don’t know

  • How the courts will ultimately rule on whether AI training constitutes 'fair use' under U.S. copyright law.
  • Whether the plaintiffs will successfully prove that Microsoft actively induced copyright infringement.
  • How a ruling against the AI companies would impact the development of future, more advanced models.

Sources

Source coverage

3 outlets

3 viewpoints surfaced

News Publishers 40%AI Developers 40%Legal Analysts 20%
  1. [1]SFGateAI Developers

    Newspaper publishers across the country are suing OpenAI and Microsoft

    Read on SFGate
  2. [2]Ars TechnicaAI Developers

    OpenAI wants its new tool to do your work for you and with you

    Read on Ars Technica
  3. [3]MediaPostNews Publishers

    'New York Times' Alters Complaint Against OpenAI And Microsoft

    Read on MediaPost

Comments

Stay informed

Every angle. Every day.

Get Artificial Intelligence stories with full source coverage and perspective breakdowns delivered to your inbox.