AI CopyrightLegal BattleJun 30, 2026, 11:31 PM· 5 min read· #5 of 5 in ai

Major Publishers and The New York Times Intensify Legal Battle Over AI Training Data

Nearly 400 local newspapers have filed a sweeping copyright lawsuit against OpenAI and Microsoft, while The New York Times amended its ongoing case to directly target Microsoft's supercomputing infrastructure.

By Factlen Editorial Team

News Publishers 40%AI Developers 40%Legal Analysts 20%
News Publishers
Argue that unauthorized AI training is mass digital theft that threatens the economic survival of journalism.
AI Developers
Maintain that training models on public data is protected fair use and critical for technological advancement.
Legal Analysts
View the lawsuits as a necessary stress test to adapt 20th-century copyright law to the generative AI era.

What's not represented

  • · Open-Source AI Developers
  • · Independent Content Creators and Artists

Why this matters

These parallel lawsuits will likely determine the economic foundation of the generative AI industry. If courts rule that AI developers must license all training data, it could force a massive transfer of wealth from tech companies to publishers and fundamentally alter how future AI models are built.

Key points

  • A coalition of nearly 400 local newspapers filed a copyright infringement lawsuit against OpenAI and Microsoft.
  • The publishers allege the tech companies scraped paywalled articles and stripped out copyright management data.
  • The New York Times amended its 2023 lawsuit to directly target Microsoft's role in building AI infrastructure.
  • The Times claims Microsoft built a custom supercomputer specifically to ingest copyrighted works at scale.
  • AI companies defend their data scraping practices as transformative fair use essential for technological progress.
  • The rulings could force AI developers to license training data or fundamentally alter how models are built.
400
Local newspapers in new lawsuit
285,000
CPU cores in Microsoft AI supercomputer
10,000
GPUs in Microsoft AI supercomputer

The legal foundation of the generative artificial intelligence industry is facing its most coordinated and aggressive challenge to date. In a span of just 24 hours, a coalition of nearly 400 local newspaper publishers launched a sweeping federal lawsuit against OpenAI and Microsoft, while The New York Times simultaneously escalated its own ongoing copyright infringement case against the two tech giants. These parallel legal maneuvers, filed in the U.S. District Court for the Southern District of New York, signal a strategic shift in how the media industry is battling unauthorized AI training, moving from broad complaints about model outputs to highly technical attacks on the underlying infrastructure.[3]

The new class-action lawsuit is led by the Local News Copyright Alliance, a group representing independent and locally owned publishing companies across 33 states, including properties like the Tahoe Daily Tribune and the Arkansas Democrat-Gazette. The coalition alleges that OpenAI and Microsoft systematically bypassed paywalls and digital access restrictions to copy hundreds of thousands of articles. According to the complaint, the tech companies used this vast repository of local journalism to train models like ChatGPT and Microsoft Copilot without ever seeking permission or offering compensation to the original creators.[1]

The local publishers argue that this unauthorized scraping constitutes a 'death knell' for an already fragile local news industry. By feeding expensive, labor-intensive local reporting into chatbots that summarize the news directly for users, the publishers claim the tech companies are actively cannibalizing the subscription and advertising revenues required to fund original journalism. The lawsuit emphasizes that local news remains one of the most trusted sources of information in America, and that its unchecked exploitation by trillion-dollar tech companies threatens the civic fabric of communities nationwide.[1]

Crucially, the local publishers' lawsuit introduces a potent new legal argument under the Digital Millennium Copyright Act (DMCA). The plaintiffs allege that when OpenAI utilized automated text extraction tools—internally referred to as 'Dragnet' and 'Newspaper'—it deliberately stripped out vital copyright management information. This allegedly included author bylines, copyright notices, and terms of use. If proven in court, this DMCA violation carries separate and significant statutory penalties that go beyond standard copyright infringement, potentially exposing the tech companies to massive financial liabilities regardless of their fair use defenses.

The legal challenges highlight both the volume of publishers involved and the massive computing infrastructure required for AI training.
The legal challenges highlight both the volume of publishers involved and the massive computing infrastructure required for AI training.

Simultaneously, The New York Times filed a heavily redacted motion to amend its December 2023 lawsuit, pivoting its legal crosshairs directly onto Microsoft. Previously viewed primarily as OpenAI's passive cloud provider and financial backer, Microsoft is now accused by the Times of actively inducing and facilitating mass copyright infringement. The amended complaint attempts to shift Microsoft's role from a neutral landlord providing generic server space to an active participant that knowingly engineered the machinery required to copy the world's news at an unprecedented scale.[3]

Simultaneously, The New York Times filed a heavily redacted motion to amend its December 2023 lawsuit, pivoting its legal crosshairs directly onto Microsoft.

The Times adapted its legal strategy following a recent Supreme Court ruling involving Cox Communications, which established a significantly tighter standard for 'contributory infringement.' Under the new legal precedent, plaintiffs can no longer simply prove that a party knew copyright infringement was happening on its servers. Instead, they must demonstrate that the party intentionally acted to induce the illegal conduct. Recognizing this shifted landscape, the Times moved quickly to realign its complaint to meet the higher burden of proof required to hold Microsoft accountable.[2]

To meet this higher legal bar, the Times alleges that Microsoft did not merely rent out off-the-shelf cloud servers to OpenAI. Instead, the amended complaint claims Microsoft built a bespoke, 'unusually complex' supercomputing system specifically designed to ingest copyrighted material. According to the filings, this custom infrastructure reportedly features over 285,000 CPU cores and 10,000 GPUs, representing a massive capital investment explicitly tailored to consume the entire internet and train the most capable large language models in history.[2]

The Times argues this infrastructure was curated to disproportionately feature high-quality journalism, ensuring that the resulting AI models could confidently mimic professional writing styles. By framing Microsoft as the architect of this ingestion engine, the lawsuit attempts to pierce the liability shield traditionally enjoyed by cloud infrastructure providers. If a court treats the provision of AI supercomputing infrastructure as active participation in scraping, every major cloud company hosting AI models could face profound new legal risks.[2]

The New York Times alleges Microsoft built a bespoke supercomputing system specifically designed to ingest copyrighted material at scale.
The New York Times alleges Microsoft built a bespoke supercomputing system specifically designed to ingest copyrighted material at scale.

OpenAI and Microsoft have consistently defended their data collection practices under the doctrine of 'fair use,' a cornerstone of U.S. copyright law that permits limited use of copyrighted material for transformative purposes. The companies maintain that training AI models on publicly available data is fundamentally transformative because the models learn statistical language patterns and facts, rather than simply storing and reproducing copyrighted expression verbatim. They argue this process is akin to a human student reading a library of books to learn how to write.[1][2]

OpenAI representatives have previously stated that ChatGPT is not a substitute for a news subscription and that restricting the use of public internet data would severely hinder technological innovation and national competitiveness. Following the recent filings, Microsoft characterized the Times' amended complaint as a 'last-ditch effort' to save its legal claims from unfavorable precedents set in other recent rulings, maintaining that their partnership with OpenAI operates entirely within the bounds of established copyright law.[2]

The outcome of these consolidated legal battles will likely dictate the future economics and technical architecture of artificial intelligence. If the courts ultimately side with the publishers and reject the fair use defense, AI companies could be forced to negotiate massive licensing deals with content creators across the globe. In the most extreme scenario, courts could order tech companies to disgorge previous profits or even wipe their existing models and rebuild them entirely from licensed data, fundamentally resetting the AI arms race.[2]

How we got here

  1. Late 2023

    The New York Times and several prominent authors file initial copyright infringement lawsuits against OpenAI and Microsoft.

  2. March 2025

    A U.S. District Judge allows the core copyright claims in the NYT lawsuit to proceed while dismissing some secondary claims.

  3. Early 2026

    The Supreme Court issues a ruling in Cox Communications that tightens the standard for contributory copyright infringement.

  4. June 24, 2026

    A coalition of nearly 400 local and regional newspapers files a new class-action lawsuit against OpenAI and Microsoft.

  5. June 25, 2026

    The New York Times files a motion to amend its complaint, shifting focus to Microsoft's bespoke supercomputing infrastructure.

Viewpoints in depth

News Publishers

Media organizations argue that AI companies are illegally profiting from their expensive original reporting.

Publishers contend that journalism is an expensive, labor-intensive endeavor that relies on subscriptions and advertising to survive. By scraping their archives and using that data to power chatbots that answer user queries directly, publishers argue AI companies are creating substitute products that siphon away their audience. They view the unauthorized use of their content not as 'learning,' but as mass digital theft that threatens the economic viability of the free press.

AI Developers

Tech companies argue that training models on public data is protected under fair use and essential for innovation.

AI developers maintain that large language models do not store or copy the text they are trained on; rather, they analyze the data to learn the statistical relationships between words. They argue this process is highly transformative and falls squarely under the 'fair use' doctrine of U.S. copyright law. Furthermore, they warn that forcing AI companies to license all training data would create an insurmountable barrier to entry, entrenching a few wealthy tech giants while stifling open-source innovation and national technological competitiveness.

Legal Scholars

Copyright experts note that AI training pushes existing legal frameworks into uncharted territory.

Legal analysts point out that current U.S. copyright law was not written with generative AI in mind. The courts are being asked to stretch the 19th-century concept of 'fair use' to cover the ingestion of billions of parameters by neural networks. Scholars suggest that if the courts find that AI companies violated the Digital Millennium Copyright Act by stripping out copyright metadata, it could provide a more straightforward path to liability than the complex fair use debate, potentially forcing Congress to draft new legislation specific to AI data rights.

What we don't know

  • How the courts will ultimately rule on whether AI training constitutes 'fair use' under U.S. copyright law.
  • Whether the plaintiffs will successfully prove that Microsoft actively induced copyright infringement.
  • How a ruling against the AI companies would impact the development of future, more advanced models.

Key terms

Large Language Model (LLM)
An artificial intelligence system trained on vast amounts of text data to understand and generate human-like language.
Fair Use
A legal doctrine in U.S. copyright law that permits limited use of copyrighted material without permission for purposes such as criticism, news reporting, teaching, and research.
Contributory Infringement
A legal concept where a party can be held liable for copyright infringement if they intentionally induce or encourage another party to commit the illegal act.
Digital Millennium Copyright Act (DMCA)
A 1998 U.S. law that criminalizes the circumvention of digital access controls and the removal of copyright management information.

Frequently asked

Why are the newspapers suing OpenAI and Microsoft?

The publishers allege that the tech companies scraped millions of their copyrighted articles without permission or payment to train AI models like ChatGPT, which then act as substitute products that reduce the newspapers' web traffic and revenue.

What is The New York Times changing in its lawsuit?

The Times is amending its complaint to argue that Microsoft actively encouraged copyright infringement by building a custom supercomputer specifically designed to ingest copyrighted material at a massive scale.

How do OpenAI and Microsoft defend their actions?

The companies argue that training AI on publicly available internet data is a transformative process that falls under 'fair use,' as the models learn patterns and facts rather than copying the original text verbatim.

What happens if the publishers win?

A victory for the publishers could force AI companies to negotiate expensive licensing deals for training data, pay significant damages, or potentially wipe and rebuild their existing AI models.

Sources

Source coverage

3 outlets

3 viewpoints surfaced

News Publishers 40%AI Developers 40%Legal Analysts 20%
  1. [1]SFGateAI Developers

    Newspaper publishers across the country are suing OpenAI and Microsoft

    Read on SFGate
  2. [2]Ars TechnicaAI Developers

    OpenAI wants its new tool to do your work for you and with you

    Read on Ars Technica
  3. [3]MediaPostNews Publishers

    'New York Times' Alters Complaint Against OpenAI And Microsoft

    Read on MediaPost
Stay informed

Every angle. Every day.

Get ai stories with full source coverage and perspective breakdowns delivered to your inbox.