Technology 15 sources · today Latest coverage 18 Sept 2026, 6:20 pm UTC

Microsoft and OpenAI Staff Raised Alarms Over AI Scraping, Court Filings Show

Newly unsealed court documents reveal internal concerns at Microsoft and OpenAI about the ethics and risks of using news content for AI training, as legal and industry scrutiny intensifies.

By Fatima Al-Rashid · First published 18 Sept 2026

In brief

  1. Court documents reveal Microsoft and OpenAI staff feared AI scraping could devastate the publishing industry and violate labor rights.
  2. Microsoft executives privately described the use of paywalled news content for AI training as the largest theft of labor in human history.
  3. OpenAI and Microsoft both acknowledged internally that their AI models posed an existential threat to news publishers' business models.
  4. These revelations surfaced as part of a copyright lawsuit brought by the New York Times against both companies in the United States.
  5. A trial is not expected before 2027, leaving the tech and publishing industries waiting for legal clarity on AI data practices.
Microsoft and OpenAI Staff Raised Alarms Over AI Scraping, Court Filings Show
Source: TechCrunch

Timeline · 7 moments

7 moments Open the full timeline →

OpenAI staff warned of existential threat to publishers

Financial Times ↗

Microsoft called AI scraping the largest theft of labor

TechCrunch ↗

Executives discussed AI undermining their own content sources

The Wall Street Journal Tech ↗

New York Times lawsuit filings reveal internal company worries

NYT > Technology ↗

NYT alleges Microsoft, OpenAI knew using news content was theft

World News CNA ↗

Judge says no trial before 2027 in copyright suit

The Hindu ↗

Microsoft and OpenAI staff feared existential threat to publishers

New York Post ↗

How it started

The roots of this controversy go back to the rapid development of large language models by companies like OpenAI and Microsoft. To build these AI systems, vast amounts of text data were needed, often sourced from news articles and other copyrighted online material.

Concerns about the ethics of using such content without permission surfaced internally at both companies. According to newly unsealed legal filings, some Microsoft executives even described these data practices as the largest theft of labor in human history. At the same time, OpenAI leaders were aware of the risks, recognizing that their AI tools could undermine the very publishers whose work powered their systems.

How it unfolded

On September 17, 2026, reports emerged that OpenAI staff and leaders were acutely aware of the existential threat their technology posed to publishers. Legal documents from the New York Times lawsuit detailed how OpenAI's business ambitions were deeply tied to the use of copyrighted content.

The same day, newly unsealed court filings revealed Microsoft executives had called the company's and OpenAI's data scraping practices theft. Internal discussions at Microsoft warned that these actions could gut the publishing industry, especially as both firms used paywalled Times content to build their AI datasets.

Further reporting confirmed that employees at both companies openly discussed the potential damage to the publishing sector. Some described the situation as a "doom loop," where the AI models could destroy the online sources they relied on.

On September 18, more details from the legal filings became public. The New York Times accused OpenAI and Microsoft of knowingly using millions of articles, mostly from the Times, to train their AI models. These revelations highlighted the unusual situation where the end-product, AI tools, could threaten the economic foundations of its content suppliers.

Where it stands

The legal battle is ongoing, with a U.S. District Judge indicating that the case will not go to trial before 2027. Meanwhile, the revelations have intensified debate in the tech and publishing sectors over the ethics and legality of AI data practices.

Both Microsoft and OpenAI now face increased scrutiny, not only from the courts but also from the public and industry partners. The outcome of this case could have far-reaching implications for how AI companies source and use data in the future.

What to watch

The main question is how the courts will rule on the legality of using copyrighted news content for AI training. Until a verdict is reached, publishers and tech companies alike are in a holding pattern, waiting to see if new legal or regulatory standards will emerge.

Written from 15 outlets' coverage of this story. Every timeline entry links to the original report.

More in Technology

All →