OpenAI lawsuit: Microsoft's 'theft of labor' memo in the NYT case
In January 2023 a Microsoft director wrote an internal memo calling AI scraping "the largest theft of labor in human history". On September 17 a federal court unsealed the summary-judgment filings in the New York Times'
In January 2023 a Microsoft director wrote an internal memo calling AI scraping "the largest theft of labor in human history". On September 17 a federal court unsealed the summary-judgment filings in the New York Times' lawsuit against OpenAI and Microsoft, and that line is now public, next to quotes from Satya Nadella, Greg Brockman and ChatGPT's own product lead. If your team trains on data it didn't license, or writes candid memos about it, this OpenAI lawsuit is the one to read.
TL;DR
- Brent Hecht, a Microsoft director of Applied Science, called AI scraping "an astonishing theft of unprecedented proportions" in a January 2023 memo. A year later his slides said Copilot's answer engine cut click-throughs to nytimes.com by up to 93 %.
- Per the Times' brief, Nadella testified that "anything that is paywalled should be licensed", and Brockman answered a "hack to get around nytimes paywall" with "ah nice."
- The brief counts more than 91,692 copies of the plaintiffs' articles in OpenAI's mid-training datasets, and over 2 million nytimes.com documents in one Common Crawl-derived set.
- Every quote comes from the Times' brief. The exhibits are still sealed, so these are the plaintiffs' picks, without context. OpenAI and Microsoft did not comment.
- Also in the episode: OpenAI's Astra for Law, a coding agent that uploads your
.gitfolder, and a $6,500 bounty for OpenAI's internal repos.
What is the NYT v. OpenAI lawsuit?
The New York Times sued OpenAI and Microsoft in December 2023, so the case turns three this December. The Daily News and the Center for Investigative Reporting are co-plaintiffs. The claim is copyright infringement: that the companies copied the papers' journalism to train models that now compete with it.
The new documents are summary-judgment motions. At this stage each side asks the judge to decide some questions without a trial, because, it argues, the facts are not in real dispute. To make that argument the plaintiffs quote the other side's own emails, slides and depositions. That is why the quotes are so pointed, and also why they need care: a brief is advocacy. TechCrunch's write-up, which everything below relies on, notes the quotes come from the Times' brief while the exhibits stay sealed, "presented without their original context".
Jason Kint, CEO of Digital Content Next, posted the filing with its redactions marked in pink:
Microsoft's "theft of labor" memo
Brent Hecht is a director of Applied Science at Microsoft. According to the brief, his January 2023 memo called AI scraping "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history".
A year later, in January 2024, he came back with a presentation. It found that Copilot's "answer engine" cut click-through rates to nytimes.com by as much as 93 % compared with traditional Bing search. That is the difference between a search result that sends you to the article and an answer that replaces it.
Hecht called it a "doom loop" that would "hurt the performance of our models and the entire web at the same time". If answer engines starve publishers of traffic, publishers stop producing, and the models have nothing new to learn from. His summary: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers." Another Microsoft document in the brief names a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."
What Nadella, Brockman and Turley said, per the brief
Satya Nadella, in a deposition this year: "anything that is paywalled should be licensed by anyone who wants to use it⦠for grounding or training". Had he known OpenAI scraped paywalled content, he said he would have "invoked [Microsoft's right to] require OpenAI to retrain its models". The same filing says "OpenAI delivered the entire GPT-3 training dataset to Microsoft", and that Microsoft sent data the other way through projects called "Project Taxi" and "Project Mango", including material from the Bing index. The brief's point: Microsoft held the dataset whose contents Nadella says he did not know.
Nick Turley, head of ChatGPT, wrote internally that publishers face an "existential threat" and that ChatGPT is "largely substitutive" and "will get more and more substitutive as they get better". Substitution is the Times' core argument: that the product replaces the paper rather than pointing to it.
Greg Brockman, OpenAI's president, said the models are "excellent at news". When researcher Nick Ryder told him about a "hack to get around nytimes paywall", Brockman replied: "ah nice."
And one engineering detail: researchers stripped copyright notices from training data because they "wouldn't want model outputting" "copyright notices" to users. Removing the notice keeps it out of the output. In a copyright case it also reads like filing the serial number off a bicycle.
How much New York Times content was in the training data?
The brief's counts, as TechCrunch reports them:
| Dataset | What the brief says it contains |
|---|---|
| OpenAI mid-training datasets | more than 91,692 copies of works by the NYT, Daily News and CIR |
| One Common Crawl-derived set | more than 2 million documents from nytimes.com |
| "Project Mango" | at least 160,903 unique works from the plaintiffs |
Common Crawl is a public archive of crawled web pages. Two million pages from one domain in one derived set shows how much of a big publisher's site ends up in "the web" unless someone filters it out.
Is AI training fair use? What the courts have said
This is the part the headline skips. Per TechCrunch, judges have "been largely favorable" to the fair-use defense so far, and earlier this month the Trump administration filed a brief defending OpenAI's unlicensed training as fair use.
The clearest ruling is Bartz v. Anthropic. In June 2025 Judge Alsup held that training on lawfully acquired books was fair use, while the roughly 7 million pirated copies Anthropic had downloaded were not. Anthropic settled in September 2025 for $1.5 billion, about $3,000 per book, with final approval in July 2026.
So the legal question is open, and on the current record it may go OpenAI's way on training itself. What the filings settle is something else. Whether the people building these systems thought it was fine, as of September 17, has an answer in their own words. Ed Newton-Rex called it possibly "the biggest admission in any of the 100+ AI copyright lawsuits". On Hacker News (over 600 points, about 500 comments) not everyone agreed it was theft: American87 remembered TechCrunch "making the argument that IP infringement != theft in the music piracy era", and Neil44 noted "the original has not gone anywhere".
A week later the authors' side of the same multidistrict case unsealed its own brief, about pirated books from LibGen, and I wrote it up in LibGen and the OpenAI lawsuit: what the unsealed authors' brief says.
What developers should take from the OpenAI lawsuit
I'm not a lawyer. These are engineering notes.
- Write the memo you'd want read aloud. Every quote above was internal. In litigation, discovery turns chat and slides into exhibits.
- Provenance and use are separate questions. Bartz split "how you got the data" from "what you did with it". Record the first for every dataset.
- Stripping metadata is a decision. Removing copyright notices to clean outputs is a normal-looking preprocessing step. In a brief it becomes intent.
- Measure what your product does to its sources. Hecht did, and found up to 93 % fewer clicks. That number is now evidence.
Also today: Astra for Law, ZCode, Hacktron
Astra for Law. The same Thursday OpenAI launched Astra for Law (HN): GPT-6 Astra, the model I stamped REVERT a week earlier, plus a legal search index covering more than 230 million URLs, with partners Harvey and Legora and case law from the Free Law Project's CourtListener. On Vals AI's Legal Research Bench it passed 54.0 % of 200 private questions, up from 38.7 % for Astra with plain web search. OpenAI calls that "a 40% relative improvement". A lawyer would call it wrong nearly half the time. cbg0 on HN: "No word on model hallucinations in the blog post".
ZCode uploads your .git. ferstar found that ZCode, Z.ai's closed-source desktop coding agent for its GLM models, packs the whole workspace, including full git history, LFS cache and reflogs, encrypts it and uploads it to Alibaba Cloud storage, on login and before every prompt. His capture: a 313 MB archive, 42,411 files, 86.6 % of it .git. The archive is encrypted with a public key from the server; only Z.ai holds the private key. His line: "A key that only the server can use serves exactly one purpose: making sure the server can read your code whenever it wants." The settings toggles did not stop it (tokenstead, HN).
Into OpenAI's repos through a picture. Hacktron (HN) used a heap overflow in libheif, the HEIF image library, to get code execution on community.openai.com, then a single sign-on misconfiguration and a GitHub integration to reach OpenAI's internal repositories, in under 72 hours. The exploit was written by Claude: Opus 4.8 could not get past address randomisation, Opus 5 produced a working exploit within three hours of its release. The whole campaign across several companies cost under $3,000 in tokens. OpenAI fixed it in about 14 hours and paid $6,500.
Verdict: NEEDS REVIEW
I stamped it NEEDS REVIEW. The quotes are real, but they are the prosecution's picks from exhibits that are still sealed, and the one judge who has ruled on this so far says training is fair use and piracy is not. The court decides the law. The memo already decided the ethics.
FAQ
What is the "largest theft of labor in human history" quote?
Per the Times' brief, Microsoft's Brent Hecht used it in a January 2023 internal memo about AI scraping.
Is the New York Times winning its lawsuit against OpenAI?
No ruling yet. The summary-judgment motions are filed, and so far courts have mostly favoured the fair-use defense on training.
Did Microsoft know OpenAI trained on New York Times articles?
The brief says OpenAI delivered the entire GPT-3 training dataset to Microsoft. Nadella testified he would have required retraining had he known paywalled content was used. The court has not decided what Microsoft knew.
Is AI training on copyrighted data legal?
In Bartz v. Anthropic, training on lawfully bought books was ruled fair use and pirated copies were not. Other cases, including this one, are still open.
Sources
- TechCrunch, the unredacted filings: https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/
- Hacker News discussion: https://news.ycombinator.com/item?id=49752056
- Jason Kint on X: https://x.com/jason_kint/status/2100611139906228629
- Ed Newton-Rex on X: https://x.com/ednewtonrex/status/2100649215290466631
- Wikipedia, Anthropic legal issues: https://en.wikipedia.org/wiki/Anthropic#Legal_issues
- OpenAI, Astra for Law: https://openai.com/index/astra-for-law/
- Astra for Law on Hacker News: https://news.ycombinator.com/item?id=49745940
- ferstar, ZCode silent workspace snapshot: https://blog.ferstar.org/en/posts/zcode-silent-workspace-snapshot/
- tokenstead write-up: https://tokenstead.ai/guides/zcode-silent-git-history-upload
- ZCode on Hacker News: https://news.ycombinator.com/item?id=49752422
- ferstar on X: https://x.com/ferstar_org/status/2100805861002355154
- Hacktron, hacking OpenAI: https://www.hacktron.ai/blog/hacking-openai
- Hacktron on Hacker News: https://news.ycombinator.com/item?id=49749656
This article expands on an episode of **The Daily Diff, a five-minute daily video on what shipped and what broke in tech.
Watch the episode Β· Subscribe on YouTube Β· the written diff lands in your inbox every morning at thedailydiff.dev.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.


