Skip to main content
ExplainerAI Copyright LawTrade-Off AnalysisAug 29, 2026, 11:57 PM· 6 min read· in opinion

Is the US 'Fair Use' Doctrine a Sufficient Legal Framework for the Generative AI Revolution?

As over 50 lawsuits test the boundaries of copyright law, the technology industry and creative economy are clashing over whether the traditional fair use doctrine can adequately govern the mass ingestion of data required to train generative AI models.

By Salma Barakat

Generative AI Developers 40%Creators and Copyright Advocates 40%Legal and Regulatory Authorities 20%
Generative AI Developers
Argues that model training is highly transformative and that requiring licenses would stifle technological innovation.
Creators and Copyright Advocates
Maintains that AI models algorithmically encode copyrighted works to create direct market substitutes without compensation.
Legal and Regulatory Authorities
Focuses on applying existing statutory frameworks to balance innovation with the protection of intellectual property rights.

The generative artificial intelligence revolution is built on a foundational legal gamble: that ingesting billions of copyrighted works to train foundation models is protected by the United States' "fair use" doctrine. For technology companies, this mass ingestion is a highly transformative extraction of uncopyrightable patterns, akin to a human reading a vast library to learn how to write. For authors, artists, and publishers, it represents the largest unauthorized reproduction of intellectual property in history, designed to create automated machines that directly compete with human creators. The question is no longer just theoretical. With over 50 lawsuits currently moving through federal courts, the U.S. legal system is being forced to decide if a flexible, 19th-century copyright doctrine can adequately govern a 21st-century automated economy. The stakes are existential for both sides: a ruling against fair use could throttle a transformative technology, while a blanket endorsement could permanently sever the economic incentives for human creation.[2][6]

The U.S. Copyright Office recently weighed in on this tension, releasing Part 3 of its comprehensive report on artificial intelligence in May 2025. The Office concluded that creating and deploying a generative AI system using copyright-protected material involves multiple acts that, absent a license or other defense, constitute prima facie infringement. Crucially, the report noted that even if the training data is discarded after the model is trained, the initial infringement analysis remains unchanged. This finding strikes at the heart of the AI industry's primary defense, shifting the legal burden entirely onto the four-factor fair use test to excuse the initial copying. By confirming that the ingestion phase is an infringing act by default, the Copyright Office has signaled that AI developers cannot simply bypass copyright law by claiming their models do not store exact copies of the input data.[1][3]

To survive this prima facie infringement, AI developers rely heavily on the first factor of the fair use test: whether the use is "transformative." In June 2025, two U.S. District Court judges in the Northern District of California issued pivotal decisions that initially seemed to validate this defense. In cases involving major AI developers, the courts recognized that using copyrighted works to train large language models is "highly transformative." The judges reasoned that the models are not reproducing the works for their intrinsic expressive value, but rather analyzing them to build statistical relationships between linguistic tokens. This interpretation aligns with earlier technology-driven fair use precedents, such as the Google Books case, suggesting that extracting uncopyrightable facts and patterns from copyrighted text is a protected activity that serves the broader public interest in innovation.[2]

A comparison of the two primary legal frameworks proposed for governing generative AI training data.

However, legal scholars argue this "transformative" narrative relies on an illusory anthropomorphism of how artificial intelligence actually functions. As detailed in a 2025 Vanderbilt Law Review article by Jacqueline C. Charlesworth, AI models do not "learn" or "reason" independently of their training data in any human sense. Instead, the copyrighted works are algorithmically encoded, token by token, into the model's vectors. Because the model's output is entirely a function of these copied materials, the argument that the training works are merely "jettisoned" after processing ignores the mechanical reality of how the copyrighted expression is stored and exploited. When a model generates output, it is relying directly on the encoded representations of the training works, undermining the claim that the original expression has been entirely left behind.[4]

However, legal scholars argue this "transformative" narrative relies on an illusory anthropomorphism of how artificial intelligence actually functions.

The most significant hurdle for the fair use defense is the fourth statutory factor: the effect on the potential market for the original work. The recent Northern District of California decisions underscored that even if training is deemed transformative, market harm remains a pivotal and potentially fatal element. Generative AI systems are increasingly used to produce outputs that directly displace the market for the ingested works. Whether by generating competing software code, writing in the specific style of a syndicated author, or utilizing retrieval-augmented generation to serve up exact excerpts from news articles, these models are acting as market substitutes. If an AI system can fulfill the consumer demand that would otherwise go to the original creator, the fair use defense traditionally collapses, regardless of how technologically innovative the copying process might be.[2][4]

Courts are also beginning to draw sharp legal distinctions based on the provenance of the training data, adding another layer of complexity to the fair use analysis. In the recent Anthropic case, the court distinguished between lawfully acquired copies of works and pirated digital libraries. The judge denied summary judgment for the AI developer regarding the use of unauthorized, pirated copies, noting that such use displaces demand for the authors' books and threatens the entire publishing market. This introduces a critical vulnerability for foundation models trained on massive, unvetted internet scrapes. If fair use cannot excuse the ingestion of unlawfully acquired data, AI developers face massive liability for the contents of open-source datasets like Books3, forcing a retroactive reckoning over how the current generation of frontier models was built.[2][6]

The mechanical reality of AI training involves algorithmic encoding, which complicates claims that training data is simply discarded.

If fair use is ultimately deemed an insufficient framework, the alternative is a mandatory licensing regime. Proponents of licensing argue it is the only way to ensure creators are compensated for the value they provide to commercial AI systems, proposing solutions like extended collective licensing. However, AI advocates warn that requiring explicit licenses for the sheer volume and diversity of content necessary to train frontier models—often trillions of tokens—is practically impossible. They caution that a strict licensing mandate would stifle innovation, entrench incumbent tech giants who possess the capital to buy data access, and severely damage U.S. competitiveness against nations with more permissive text-and-data-mining exceptions. Ultimately, the U.S. is attempting to solve a macroeconomic structural shift using a microeconomic legal tool, forcing courts to decide whether fair use can stretch to cover the wholesale transfer of value from the creative economy to the technology sector.[1][5]

As the legal landscape fractures, the technology industry is quietly preparing for a hybrid future where fair use and licensing coexist. Major AI developers have already begun signing multi-million-dollar licensing agreements with large publishers and stock image repositories, effectively hedging their bets against an adverse Supreme Court ruling. This dual-track approach suggests that while companies will continue to aggressively litigate their fair use rights for broad internet scraping, they acknowledge that premium, high-fidelity data requires a commercial transaction. The ultimate resolution will likely require congressional intervention, as the current strategy of litigating 50 separate cases through various appellate circuits promises years of uncertainty. Until a unified legal framework emerges, the generative AI revolution will continue to operate under a cloud of copyright liability, balancing unprecedented technological capability against the unresolved property rights of the human creators who fueled it.[3][7]

Competing readings

Relying on the Traditional Fair Use Doctrine

Maintaining the status quo where AI companies ingest data without licenses, defending the practice under the four-factor fair use test.

**For:** Preserves rapid technological innovation and U.S. global competitiveness by allowing developers to train on the massive, diverse datasets required for frontier models without prohibitive transaction costs. Prevents the entrenchment of a few wealthy tech monopolies by keeping training data accessible to open-source developers and startups. **Against:** Forces individual creators into asymmetric, expensive litigation to protect their work. Fails to compensate the human labor that creates the high-quality data AI relies on, potentially collapsing the economic incentives for future original creation. **Evidence:** In June 2025, two U.S. District Courts in California ruled that training LLMs is "highly transformative" under the first fair use factor, supporting the doctrine's applicability. However, over 50 pending lawsuits demonstrate the friction this approach causes. **Fits well when:** The AI model is trained strictly on lawfully acquired data and generates purely abstract, non-competing outputs (e.g., medical research analysis or backend coding). **Does not fit when:** The model is trained on pirated databases (like Books3) or acts as a direct market substitute for the ingested works, such as generating stylistic art or utilizing retrieval-augmented generation (RAG) to bypass original publishers.

Implementing a Mandatory Licensing Framework

Requiring AI developers to obtain explicit permission and pay royalties for any copyrighted material used in training datasets.

**For:** Re-establishes property rights and ensures creators are financially compensated for the value their work provides to multi-billion-dollar commercial AI systems. Creates a transparent market for high-quality training data, reducing the legal risks of copyright infringement for enterprise AI users. **Against:** The sheer volume of data required (trillions of tokens) makes individual licensing practically impossible. It would likely result in a fragmented AI ecosystem where only the largest tech conglomerates can afford to build foundation models, stifling open-source development. **Evidence:** The U.S. Copyright Office's May 2025 report concluded that using copyrighted material for AI training is prima facie infringement absent a defense, validating the underlying property right. Meanwhile, some nations are exploring extended collective licensing to solve the transaction-cost problem. **Fits well when:** Applied to highly specific, premium datasets (like news archives, proprietary codebases, or specialized medical journals) where the data owners are easily identifiable and the value extraction is direct. **Does not fit when:** Applied retroactively to the broad, public internet scraping required to teach a foundation model basic language syntax and general world knowledge, where clearing billions of micro-licenses is mathematically unfeasible.

4
Statutory fair use factors
50+
Pending AI copyright lawsuits
2
Recent NDCA fair use rulings

Sources

Source coverage

7 outlets

3 viewpoints surfaced

Generative AI Developers 40%Creators and Copyright Advocates 40%Legal and Regulatory Authorities 20%
  1. [1]U.S. Copyright OfficeLegal and Regulatory Authorities

    Copyright and Artificial Intelligence, Part 3: Generative AI Training

    Read on U.S. Copyright Office
  2. [2]Jones DayGenerative AI Developers

    Two U.S. Courts Address Fair Use in Generative AI Training Cases

    Read on Jones Day
  3. [3]Wiley ReinLegal and Regulatory Authorities

    Copyright Office Issues Key Guidance on Fair Use in Generative AI Training

    Read on Wiley Rein
  4. [4]Scholarship@Vanderbilt LawCreators and Copyright Advocates

    Generative Al's Illusory Case for Fair Use

    Read on Scholarship@Vanderbilt Law
  5. [5]Copyright AllianceCreators and Copyright Advocates

    5 Takeaways from the Copyright Office's Report on Generative AI Training

    Read on Copyright Alliance
  6. [6]The AtlanticCreators and Copyright Advocates

    The Pirated-Book Database That Built AI

    Read on The Atlantic
  7. [7]Factlen Editorial TeamLegal and Regulatory Authorities

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get opinion stories with full source coverage and perspective breakdowns delivered to your inbox.