How AI Crawlers Actually Read the Web, and Why Publishers Are Demanding Model Destruction
The Seattle Times and Newsday have sued OpenAI and Microsoft, escalating the copyright fight by demanding the destruction of AI models trained on their data. The conflict highlights the mechanics of AI web crawlers and the limitations of the voluntary opt-out systems that govern them.
By Lila Morgan
- News Publishers
- Argue that scraping paywalled journalism without compensation is copyright infringement that destroys their business model.
- AI Developers
- Argue that model training is a transformative fair use of publicly accessible data that drives technological progress.
- Open Web Advocates
- Concerned that the scraping arms race will force all useful information behind strict paywalls, breaking the open internet.
Perspectives this story doesn't cover
- Independent content creators whose work is scraped but who lack the resources to sue
- Academic researchers relying on web scraping for non-commercial studies
Website administrators decide exactly which artificial intelligence bots can read their pages by editing a single text file, a choice they can update at any time. But for data scraped before those rules were written, the U.S. District Court for the Southern District of New York will now decide whether the resulting AI models must be dismantled entirely. On September 4, 2026, The Seattle Times and Newsday filed a federal lawsuit against OpenAI and Microsoft, escalating the ongoing conflict over digital copyright.[1][2]
The two publications accuse the technology companies of scraping their paywalled journalism without permission to train generative AI systems, including ChatGPT and Microsoft Copilot. Unlike previous lawsuits that primarily sought financial compensation, this filing demands the physical destruction of the training datasets and the AI models that incorporate their work.[1][2]
The legal confrontation exposes the underlying mechanics of how artificial intelligence systems actually acquire their knowledge. To build a large language model, developers deploy automated software programs known as web crawlers or bots. These programs systematically navigate the internet, following hyperlinks from one page to the next, downloading the text they encounter along the way.[3]
Web crawlers have long served as the foundation of digital discovery. Search engines like Google and Bing use them to index the web, allowing users to find relevant pages. When a search crawler visits a site, the implicit transaction is an exchange of access for visibility: the search engine reads the content, and in return, it directs human traffic back to the publisher.[3]
AI training crawlers operate under a different economic model. When a bot like OpenAI's GPTBot reads a news article, it is not indexing the page to send a reader there later. Instead, it is extracting the text to teach a neural network how to generate human-like language, predict the next word in a sequence, and answer user queries directly.[3]
The publishers argue this extraction fundamentally undermines their business. The lawsuit alleges that the AI products can reproduce passages from their reporting and closely paraphrase full articles, providing users with answers that reduce the need to visit the original website or purchase a subscription.[1][2]
The publishers argue this extraction fundamentally undermines their business.
To control these automated visitors, the internet relies on a voluntary protocol established in 1994: the robots.txt file. This simple text document sits at the root directory of a website and acts as a traffic controller, listing specific instructions for different bots.[3]
An administrator can add a line of code to their robots.txt file explicitly disallowing a specific user-agent, such as GPTBot. When a compliant crawler arrives at the site, it reads the file, sees the restriction, and leaves without downloading the content. When a site administrator updates this file, it typically takes roughly 24 hours for major AI systems to adjust their crawling behavior.[3]
However, the robots.txt system has significant limitations. It is entirely an honor system; malicious or poorly configured bots can simply ignore the instructions and scrape the site anyway. More importantly for the current litigation, a robots.txt directive only governs future behavior. It cannot retroactively delete data that a bot scraped before the rule was implemented, nor can it extract that data from a model that has already been trained.[3]
The distinction between different types of AI crawlers has also complicated the landscape for publishers. Major AI laboratories now operate multiple, independently governed bots. For example, OpenAI deploys GPTBot specifically to collect training data, but it uses a separate crawler, OAI-SearchBot, to index pages for its real-time search features.[3]
If a publisher blocks GPTBot, they prevent their future articles from being used to train the next generation of models. But if they accidentally block OAI-SearchBot in the process, they remove themselves entirely from ChatGPT's search answers, cutting off a growing source of referral traffic. Industry data cited in similar publisher lawsuits indicates that search referral traffic to midsize publishers fell by 47 percent year-over-year in December 2025.[3]
The Seattle Times and Newsday are not the first to challenge this ecosystem. The New York Times filed a similar lawsuit in December 2023, and a coalition of nearly 400 local newspapers followed suit in June 2026. The addition of two prominent regional papers underscores the specific vulnerability of local newsrooms, which rely heavily on subscription revenue to fund their reporting.[1][2]
Microsoft has publicly responded to the latest filing with a measured stance. A spokesperson for the company stated, "While we're surprised by the lawsuit, we appreciate the importance of local journalism and we're always happy to sit down and explore solutions to this type of dispute."[1][2]
The core legal question now rests on the doctrine of fair use. AI companies maintain that training models on publicly available data is a transformative act protected by copyright law, while publishers argue it is wholesale theft that directly competes with their original products. The court's eventual ruling will establish the boundaries of what automated systems are permitted to learn from the open web.[1][3]
What to know
- The Seattle Times and Newsday filed a federal lawsuit against OpenAI and Microsoft on September 4, 2026.
- The publishers are demanding the destruction of AI models trained on their copyrighted journalism.
- AI companies use automated web crawlers to extract text from the internet to train large language models.
- Website owners can block these crawlers using a robots.txt file, but the protocol is a voluntary honor system.
- A robots.txt directive only governs future crawling and cannot retroactively delete data that has already been scraped.
Key terms
- Web Crawler
- An automated software program that systematically browses the internet to read and download web page content.
- robots.txt
- A standard text file placed on a website that provides instructions telling automated bots which pages they are allowed or forbidden to visit.
- User-agent
- A string of text that a web crawler uses to identify itself to a website's server, allowing administrators to set specific rules for different bots.
- Fair Use
- A legal doctrine in U.S. copyright law that permits limited use of copyrighted material without acquiring permission from the rights holders, often cited as a defense by AI developers.
- Large Language Model (LLM)
- An artificial intelligence system trained on vast amounts of text data to understand and generate human-like language.
Reader questions
What is an AI web crawler?
An AI web crawler is an automated software program that navigates the internet to download text and data. Unlike search engine crawlers that index pages to help users find them, AI crawlers extract the data to train large language models.
How can a website block an AI crawler?
Website administrators can block specific crawlers by adding a 'Disallow' rule for that bot's user-agent (such as GPTBot) in their site's robots.txt file. However, this relies on the bot voluntarily honoring the request.
What are the Seattle Times and Newsday asking for?
In addition to unspecified financial damages, the two newspapers are demanding the physical destruction of the training datasets and the AI models that incorporate their copyrighted journalism.
Does blocking an AI crawler remove a site from AI search?
It depends on the specific bot blocked. Blocking a training crawler like GPTBot stops data collection for future models, but blocking a search crawler like OAI-SearchBot will remove the site from real-time AI search answers.
Sources
[1]TechCrunchNews PublishersSeattle Times and Newsday are the latest publications to sue OpenAI and Microsoft
Read on TechCrunch →
[2]EngadgetNews PublishersTwo more news organizations sue OpenAI and Microsoft for copyright infringement
Read on Engadget →
[3]Factlen Editorial TeamOpen Web AdvocatesSynthesis by Factlen editorial team
Read on Factlen Editorial Team →
Comments
More in Technology
See all →Spectrum Regulation
Why Bluetooth Jammers Are Illegal: The Mechanics of 2.4 GHz Interference
4 sources
Lithography Physics
The Rayleigh Criterion: How Wavelength and Numerical Aperture Actually Constrain Chip Scaling
8 sources
Smart TV Privacy
LG Smart TVs Caught Logging Audio and Scanning Local Networks in Standby
4 sources
LMR Battery Tech
LG Energy Solution and Seoul National University Resolve Gas Buildup in Cobalt-Free LMR Batteries
5 sources
Every angle. Every day.
Get Technology stories with full source coverage and perspective breakdowns delivered to your inbox.




