New York Times lawsuit reveals internal Microsoft and OpenAI documents on AI data scraping
Unredacted internal memos from Microsoft and OpenAI have been filed in the NYT-led copyright case, showing executives label massive content harvesting as unprecedented labor theft and warn of a destructive revenue loop.

In the ongoing copyright lawsuit filed by The New York Times and a coalition of other publishers against Microsoft and OpenAI, the plaintiffs have introduced a set of unredacted internal documents that illuminate how the two technology giants internally discuss the large‑scale harvesting of copyrighted material for the purpose of training generative AI models.
The court filings contain emails, internal briefing papers and analytical reports that were not censored before being entered into the public record, as reported by Ars Technica on September 17, 2026 and by Golem.de on September 18, 2026.
Executive language frames the practice as unprecedented theft
One internal memo quotes Brent Hecht, Microsoft’s director of applied sciences, describing the massive content collection as "the largest theft of labor in human history" and noting that he had internally questioned the company’s reliance on a fair‑use defense for such activities.
Hecht’s comment, as presented in the filings, is presented as a personal viewpoint rather than a formal legal analysis, a distinction that Microsoft later stressed in its public response to the allegations.
Potential economic feedback loop
A separate internal Microsoft document warns of a "destructive loop": AI products that draw on copyrighted content could erode the revenue streams of the original content producers, which in turn would diminish the pool of material available for future model training.
The memo argues that this cycle could undermine the long‑term health of the publishing ecosystem, a concern that the plaintiffs echo in their complaint to the court.
Evidence of traffic decline and user behavior
The plaintiffs have submitted data indicating click‑through‑rate drops ranging from 83 % to 93 % for certain headlines and from 51 % to 94 % for other headlines after AI‑generated answers were displayed in place of the original articles. The figures are offered as evidentiary support, not as adjudicated facts.
Internal OpenAI messages reveal that some employees believed users would be less inclined to click on a link when the answer was already supplied within the AI response, reinforcing the publishers’ claim of reduced traffic.
- Brent Hecht called the scraping "the largest theft of labor in human history"
- Microsoft warned of a destructive revenue loop for content producers
- Satya Nadella testified that AI assistants can replace a site visit by delivering information directly
- OpenAI staff expected lower click‑through rates when answers were pre‑generated
Satya Nadella, Microsoft’s chief executive, testified under oath that AI assistants could effectively replace a user’s visit to the original source by providing the requested information directly, a statement that aligns with the internal concerns about traffic loss.
The plaintiffs also allege that a paywall‑circumvention technique was discovered, enabling a third‑party vendor to supply roughly 1.8 million New York Times articles under commercial‑use restrictions that were later leveraged for model training.
Microsoft has responded that the quoted statements reflect an individual employee’s perspective and do not constitute the company’s official legal position. Both Microsoft and OpenAI continue to argue that their use of copyrighted material falls within the bounds of fair use, a claim that has yet to be adjudicated by a court.
No judicial decision on the merits of these allegations has been rendered; the documents remain part of the evidentiary record while the case proceeds toward a substantive ruling.
For English‑speaking organizations, the emergence of these internal documents signals a heightened risk profile for companies that rely on AI‑generated content. The disclosed internal doubts about fair use and the quantified traffic losses suggest that future litigation could impose stricter licensing requirements, compel more transparent data‑sourcing practices, and potentially increase the cost of deploying large‑scale language models that depend on copyrighted text.
In New York, where the lawsuit was filed, legal analysts note that the public availability of these memos may influence not only the outcome of this particular case but also the broader regulatory conversation about AI training data, prompting publishers and tech firms alike to reassess their data‑collection strategies.
Sources
- Microsoft exec called AI scraping the largest theft of labor in human historyArs Technica · September 17, 2026
- Ungeschwärzte Dokumente belasten Microsoft und OpenAIGolem.de · September 18, 2026



