Natural-Language Disclaimers Fail Article 4(3) Machine-Readability Rules: Why Plain-Text Terms Cannot Block AI Training Crawlers
Under European copyright law and the Artificial Intelligence Act, rightsholders must use specific technical protocols to legally block text and data mining. Standard terms of service fail this standard, leaving digital content exposed to commercial model training.
By Sofia Matos
In short
- Article 4(3) of the EU Copyright Directive permits commercial text and data mining unless rights are expressly reserved via machine-readable technical protocols.
- The EU AI Act mandates that all general-purpose AI providers respect these machine-readable opt-outs, extending the rule's enforcement to global developers.
- Recent court rulings confirm that natural-language terms of service fail the machine-readability test, leaving content legally exposed to automated AI training crawlers.
In this article
On September 3, 2026, the German Federal Court of Justice heard arguments on a question that dictates the future of digital copyright: whether a website's plain-text terms of service can legally block artificial intelligence training crawlers. The case tests the boundaries of European law.[1][5]
The dispute centers on Article 4(3) of the European Union's Copyright in the Digital Single Market (CDSM) Directive. The directive permits commercial text and data mining on lawfully accessible content, provided the rightsholder has not expressly reserved those rights.[1][2]
Crucially, the law specifies exactly how that reservation must be communicated online. Rightsholders must express their opt-out "in an appropriate manner, such as machine-readable means." This single phrase has reshaped how digital intellectual property is defended against automated extraction.[1][5]
For thousands of publishers and creators, this technical requirement has invalidated their legal defenses. A natural-language disclaimer buried in a website's footer or terms of service is invisible to automated crawlers, rendering it legally insufficient under the directive's strict standard.[2][5]
The Legal Architecture of Article 4(3)
The CDSM Directive, adopted in 2019, established two distinct exceptions for text and data mining. Article 3 created a mandatory, non-overrideable exception for scientific research conducted by recognized research organizations and cultural heritage institutions, shielding academic work from commercial restrictions.[1][4]
Article 4, however, governs commercial text and data mining, which includes the massive data ingestion required to train general-purpose AI models. Unlike the research exception, Article 4 operates on an opt-out basis, placing the burden of action entirely on the content creator.[4][5]
If a creator fails to properly reserve their rights, commercial AI developers are legally permitted to scrape and analyze their publicly available works. The directive explicitly notes that for online content, this reservation requires a format that automated systems can process without human intervention.[1][5]
The Hamburg Higher Regional Court affirmed this strict interpretation on December 10, 2025. In the Kneschke v. LAION decision, the court ruled that a website's natural-language ban on automated bots did not meet the machine-readable threshold at the time the scraping occurred.[1][5]
"A free-text caption is neither technically respected nor legally sufficient as a reservation of rights," notes legal analysis from Promise Legal. The ruling established that human-readable legal text cannot bind a machine that is not programmed to read it.[2][6]
The Failure of Natural Language
The disconnect stems from how AI training pipelines actually operate in practice. When a developer compiles a dataset, they deploy automated web crawlers that index millions of URLs per hour. These bots parse HTML structure, metadata, and specific protocol files to navigate the web.[2][5]
They do not, however, employ natural language processing to read and interpret a website's terms of service before downloading an image or text file. A crawler sees a "No AI Training" clause as just another string of text to be scraped, not as a binding access control.[5][6]
Consequently, standard software-as-a-service agreements and website footers declaring "all rights reserved" fail the Article 4(3) test. Because the crawler cannot automatically evaluate the legal condition, the opt-out is not machine-readable, and the commercial text and data mining exception remains legally valid.[2][5]
This creates a severe vulnerability for content creators who rely on traditional copyright notices. Updating a legal agreement to prohibit machine learning extraction provides no protection if the technical delivery mechanism remains human-readable text that the ingestion engine simply ignores.[2][6]
To successfully block ingestion, the legal intent must be translated into a technical signal. The law does not mandate a single format, but it requires a protocol that a crawler can ingest and obey programmatically at scale without requiring human legal review.[5][6]
The AI Act's Extraterritorial Reach
The stakes of this technical distinction escalated dramatically with the enforcement of the EU Artificial Intelligence Act. Since August 2, 2025, providers of general-purpose AI models have faced binding copyright compliance obligations under Article 53(1)(c) of the sweeping regulatory framework.[1][3]
The regulation requires AI providers to "put in place a policy to comply with Union copyright law, and in particular to identify and comply with, including through state of the art technologies, a reservation of rights expressed pursuant to Article 4(3)."[3][5]
This provision transforms the CDSM Directive's opt-out mechanism from a theoretical defense into an active compliance mandate. AI developers must actively scan for and respect machine-readable reservations, or face severe regulatory penalties under the AI Act's enforcement mechanisms.[1][3]
Furthermore, the AI Act applies extraterritorially. Any general-purpose AI model placed on the European Union market must comply with these copyright policies, regardless of where the actual data scraping or model training occurred, fundamentally altering global AI development practices.[2][3]
"The EU's position, articulated in Recital 106, holds that copyright obligations attach to models offered in the EU market," explains the AI Governance Desk. This forces US and Asian AI developers to build infrastructure that detects European machine-readable opt-outs.[3][6]
Machine-Readable Standards in Practice
With natural language disqualified, publishers are adopting specific technical protocols to secure their Article 4(3) rights. The most widely recognized mechanism is the robots.txt file, formatted according to the RFC 9309 standard, which instructs web crawlers on permitted access paths.[1][5]
While robots.txt was originally designed as a voluntary request for search engine indexers, the AI Act's Code of Practice explicitly recognizes it as a valid machine-readable protocol for expressing reservations, granting the text file unprecedented legal weight.[5][6]
However, robots.txt operates at the domain or directory level, making it clumsy for granular rights management. For individual assets, rightsholders are turning to the Text and Data Mining Reservation Protocol (TDMRep), developed by the World Wide Web Consortium.[2][5]
TDMRep allows publishers to host a specific JSON file that signals rights reservations programmatically. Alternatively, creators can embed HTTP headers or HTML meta tags, such as the noai directive, directly into the page code to protect specific digital assets.[2][5]
A robust defense now requires a multi-layered approach. Legal experts advise combining a robots.txt block, TDMRep metadata, and explicit terms of service to ensure the reservation is both technically effective and legally unassailable in court if a dispute arises.[1][2]
The Burden on Content Creators
The shift toward machine-readable requirements has fundamentally altered the balance of power in digital copyright. Historically, copyright protection was automatic upon creation, requiring no technical expertise or active maintenance from the author to remain legally enforceable against unauthorized commercial exploitation.[2][6]
Article 4(3) reverses this dynamic for AI training. By requiring an active, technically specific opt-out, the law places the administrative and financial burden entirely on the rightsholder. Creators must now understand web protocols, edit server configurations, and monitor evolving crawler standards.[2][6]
This technical barrier disproportionately affects independent artists, freelance photographers, and small publishers who lack dedicated IT departments. While large media conglomerates can easily deploy comprehensive robots.txt rules and TDMRep headers, individual creators often rely on portfolio platforms that restrict backend access.[2][6]
If a hosting platform does not natively support machine-readable opt-outs, the creator's work remains legally exposed to commercial extraction, regardless of their personal wishes. The legal framework effectively penalizes those without the technical means to express their rights in code.[2][6]
The Open-Weight Complication
Even when a machine-readable opt-out is perfectly executed, the architecture of generative AI introduces a structural failure point. The opt-out mechanism offers no remedy once protected works have already been ingested into a fully trained and finalized neural network.[1][5]
This deficiency is particularly acute with open-weight models. If a model trained on legally reserved data is released publicly, the rightsholder cannot compel the retraining of the deployed model or prevent third parties from using the existing weights for inference.[1][6]
Key terms
- Text and Data Mining (TDM)
- The automated analytical technique used to extract patterns, trends, and correlations from digital content, serving as the foundational process for training AI models.
- CDSM Directive
- The 2019 EU Copyright in the Digital Single Market Directive, which established the legal rules for when commercial entities can scrape online data.
- Machine-Readable
- Data presented in a structured format that automated computer systems and web crawlers can process and obey without human intervention.
- TDMRep
- The Text and Data Mining Reservation Protocol, a standardized technical method developed by the W3C for rightsholders to signal their copyright opt-outs programmatically.
- robots.txt
- A standard text file placed on a website's server that instructs automated web crawlers which pages or files they are permitted to access.
Frequently asked
Does a standard copyright symbol prevent AI training?
No. A traditional copyright symbol or 'all rights reserved' text is not machine-readable. Under Article 4(3) of the CDSM Directive, commercial AI crawlers can legally ignore natural-language notices and ingest the content.
Is a robots.txt file legally binding for AI companies?
Yes, under the EU AI Act. While robots.txt was historically just a voluntary request, the AI Act's Code of Practice explicitly recognizes it as a valid machine-readable protocol that general-purpose AI providers must respect.
Does this European law affect AI companies based in the US?
Yes. Article 53(1)(c) of the EU AI Act applies to any general-purpose AI model placed on the European market, forcing global developers to respect machine-readable opt-outs regardless of where the model was trained.
Can an opt-out remove my data from an already trained model?
No. The text and data mining exception applies to the act of copying data during the training process. A machine-readable opt-out prevents future scraping but offers no mechanism to force a deployed model to 'unlearn' previously ingested data.
Viewpoints in depth
AI Model Developers
Argue that standardized, machine-readable signals are the only technically feasible way to respect copyright at the scale of web crawling.
From an engineering perspective, AI developers maintain that reading and interpreting natural-language terms of service across millions of domains is computationally impossible. Crawlers are designed to ingest data rapidly, relying on standardized protocols like robots.txt or HTTP headers to dictate access rules programmatically. Developers argue that without a universal, machine-readable standard, they would be forced to employ human legal teams to review every website's footer before scraping, effectively halting the development of large-scale datasets. They view Article 4(3) as a necessary compromise that enables innovation while providing a clear, automated pathway for rightsholders to opt out.
Content Creators & Publishers
Argue that the machine-readability requirement places an unfair technical burden on artists to actively defend their default copyright protections.
Publishers and independent creators view the machine-readability mandate as a fundamental erosion of their intellectual property rights. Historically, copyright was automatic; a simple 'all rights reserved' notice was sufficient to deter commercial exploitation. Under the current interpretation of Article 4(3), creators are forced to become systems administrators, managing server-level configurations and JSON files just to maintain the rights they already own. Advocacy groups argue this creates a two-tiered system where only large media corporations with dedicated IT departments can successfully protect their content, leaving independent artists legally exposed simply because they lack the technical infrastructure to deploy a TDMRep tag.
Legal & Regulatory Experts
Emphasize that the law requires strict technical compliance, rendering traditional natural-language contracts ineffective against automated systems.
Legal scholars point out that the European framework was deliberately designed to separate human-readable contracts from machine-executable commands. The courts, as seen in the initial LAION ruling, are strictly interpreting the text of the CDSM Directive: if a crawler cannot read the opt-out, the opt-out does not legally exist for the purposes of that crawler. Regulatory experts note that while this creates friction for creators, the introduction of the AI Act's Article 53(1)(c) provides a powerful enforcement mechanism. By forcing AI companies to actively detect and respect these technical signals globally, the law establishes a rigid but enforceable boundary between public data and reserved intellectual property.
- AI Model Developers
- Argue that standardized, machine-readable signals are the only technically feasible way to respect copyright at the scale of web crawling.
- Content Creators & Publishers
- Argue that the machine-readability requirement places an unfair technical burden on artists to actively defend their default copyright protections.
- Legal & Regulatory Experts
- Emphasize that the law requires strict technical compliance, rendering traditional natural-language contracts ineffective against automated systems.
Perspectives this story doesn't cover
- Open-Source AI Advocates
- Web Archiving Organizations
Sources
[1]IP Global GuardLegal & Regulatory ExpertsTo opt out of AI training in the EU, a rightsholder must expressly reserve its text and data mining (TDM) rights under Article 4(3)
Read on IP Global Guard →
[2]Promise LegalContent Creators & PublishersThe EU TDM Article 4(3) reservation framework requires opt-outs in machine-readable form
Read on Promise Legal →
[3]AI Governance DeskAI Model DevelopersArticle 53(1)(c) and (d): The Dual Obligation Framework
Read on AI Governance Desk →
[4]University of CyprusAI Model DevelopersThe TDM Exceptions in the CDSM Directive Through the Lens of AI Training
Read on University of Cyprus →
[5]Black Alpaca LegalContent Creators & Publisherstxt and AI: EU Legal Situation and TDM Opt-out
Read on Black Alpaca Legal →
[6]Factlen Editorial TeamLegal & Regulatory ExpertsSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
More in Artificial Intelligence
See all →Substantial Modification
Why Enterprise Fine-Tuning Converts AI Deployers Into Legal Providers
6 sources
AI Compliance
The Five Steps of an Algorithmic Impact Assessment Regulators Use to Mandate AI Risk Mitigation
3 sources
AI Explainability
The Inverse Relationship Between AI Model Complexity and Decision Explainability
11 sources
Frontier AI
The 10^26 FLOP Threshold: How the US Government Monitors Frontier AI
4 sources
Comments
Every angle. Every day.
Get Artificial Intelligence stories with full source coverage and perspective breakdowns, free every day.




