Microsoft exec called AI scraping ‘largest theft of labor

Microsoft exec called AI scraping ‘largest theft of labor

Microsoft exec called AI scraping ‘largest theft of labor’

TL;DR: Newly unredacted court filings in The New York Times’ copyright case against OpenAI and Microsoft quote a Microsoft director of applied science calling AI scraping “the largest theft of labor in human history” — in his own company’s internal documents. The same filing pairs those words with Microsoft’s own measurement showing Copilot answer engines cut click-throughs to publisher sites by up to 94% versus normal Bing search. The admissions are aimed squarely at the fair use defence OpenAI and Microsoft are relying on, and they land just as Malaysia prepares to decide its own rules on AI training data.

Microsoft has spent three years arguing, in public, that training AI models on scraped web content is fair, transformative and lawful. Its own staff were apparently not so sure.

On 17 September 2026, a US court unsealed the public version of a 92-page summary judgment brief filed by news publishers led by The New York Times in the Southern District of New York (docketed in the consolidated proceedings, MDL No. 25-md-3143, before Judge Sidney H. Stein). Much of the material in it had been blacked out for two years. What emerged is a paper trail in which the defendants’ own employees describe AI scraping in language no lawyer would have drafted for them.

What the unredacted filings reveal about AI scraping at Microsoft

The most-quoted line comes from Brent Hecht, Microsoft’s Director of Applied Science. Writing in an internal memo in January 2023, according to the publishers’ brief, Hecht described the mass harvesting of news content for model training as “the largest theft of labor in human history” — and, in a fuller passage, “an astonishing theft of unprecedented proportions.”

He did not stop there. In the same internal record cited in the filing, Hecht noted that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.” He also wrote that if courts blessed this kind of AI scraping as fair use, it would arguably “make a complete mockery of the idea of ‘fair use.’”

Microsoft’s response has been to narrow the frame: a company spokesperson said Hecht’s comments reflected his individual views rather than Microsoft’s position, and a separate Microsoft filing reportedly framed him as someone on the payroll to play the contrarian. That characterisation runs into an obvious problem — Hecht was not a lone blogger. Microsoft’s own data and presentations are cited throughout the same filing, and the numbers inside them are the part that is hardest to argue away.

The quotes, it should be said plainly, come from the publishers’ filing rather than from the underlying exhibits, which remain partly sealed. Quotes are also presented without full original context. But the language is on the public record now, filed by lawyers who can be sanctioned for misquoting, and reformatted in every major tech and business outlet within a day of being unsealed.

OpenAI’s internal record, as cited in the same filing, is no kinder to the fair use story. Nick Turley, the executive who led ChatGPT, wrote that publishers face an “existential threat” from products like his, describing them as “largely substitutive, period” and warning they “will get more and more substitutive as they get better.” Greg Brockman, OpenAI’s co-founder and president, said the models were “excellent at news.” One OpenAI engineer put the traffic problem about as bluntly as it can be put: “no matter how prominently we show the links, users won’t click.”

The ‘doom loop’: Microsoft’s own data on Copilot referrals

The AI scraping debate usually turns abstract the moment lawyers start talking about “market substitution.” Microsoft’s own telemetry made it concrete.

An internal Microsoft presentation written by Hecht in January 2024, as described in the filing, labelled the company’s AI content strategy a “doom loop” that would “hurt the performance of our models and the entire web at the same time.” The presentation contained this line, quoted in the filing:

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

Another Microsoft line cited in the same filing is shorter still: “LLMs are a product that destroys its supply chain.”

Then there are the click-through numbers, which is where AI scraping stops being a philosophical argument and becomes arithmetic. Microsoft’s data, per the filing, compared referrals from Bing Chat and Copilot against ordinary Bing web search:

Publisher group Click-through rate drop vs Bing search
The New York Times 87%–93% (as much as 93%)
Daily News titles (Alden Global Capital) 83%–91%
Ziff Davis (CNET, IGN, PCMag, Mashable and others) 51%–94%

An earlier, redacted version of the same filing had these percentages blacked out. The publishers’ point is simple: Microsoft measured the damage, wrote it down, described it internally as a doom loop, and shipped the product anyway.

For anyone who publishes anything on the web — including Malaysian blogs, review sites, news portals and small e-commerce content shops — that table is the whole story of the last two years of declining search referrals, told by the company that owns the search engine.

The data pipeline: Project Taxi, Project Mango and paywall bypass

The filing also reconstructs how the training data allegedly moved between the two companies, using internal codenames that had not previously appeared in public documents.

  • OpenAI’s dataset to Microsoft. OpenAI delivered the entire GPT-3 training dataset to Microsoft, which used it to evaluate how to deploy OpenAI’s models inside its own commercial products. The publishers note GPT-3 was already trained by then, so the transfer had nothing to do with training a model at issue in the case.
  • Project Taxi. Over three years from 2019 to 2022, Microsoft supplied OpenAI with a copy of the Bing Index — described by Microsoft’s own counsel as a compilation of “billions” of webpages gathered for conventional search. OpenAI persuaded Microsoft to “sell” it the index for a price that remains redacted. OpenAI’s then chief research officer Bob McGrew called it “a trade” in October 2021: “the idea [was] that we’re giving data to them and they are giving data to us.” The publishers say no copyright holder was asked.
  • Project Mango. Microsoft ran a crawler on OpenAI’s behalf to “collect as many of the documents as possible for their training,” with OpenAI paying Microsoft an undisclosed sum. The resulting dataset, according to the filing, contains copies of at least 160,903 unique works belonging to the plaintiff publishers.
  • Paywall circumvention and notice stripping. The filing alleges that OpenAI staff discussed a “hack” to bypass The Times’ paywall without detection — and that Brockman replied “ah nice.” It further alleges copyright notices were removed from training copies in 99.4% to 100% of cases, and that a so-called “Giraffe” Bloom filter was built for de-duplication after the lawsuits were filed.

Scale, in the publishers’ numbers: roughly 3.9 million training copies of Times works (about 944,655 unique works) and 7.4 million copies of Daily News works (around 1.39 million unique). A Common Crawl-derived dataset alone held more than 2 million documents from nytimes.com, and the mid-training sets contained north of 91,692 copies of works from the Times, the Daily News and the Center for Investigative Reporting.

The words “AI scraping” do a lot of hiding here. What the filing describes is not a passive slurp of a public web page; it is a described pipeline in which crawling, paywall workarounds, dataset assembly, sanitisation and inter-company trading were each somebody’s job.

Why the ‘largest theft of labor’ line cuts against the fair use defence

US fair use analysis has four factors, and the one that has decided recent cases is the fourth: whether the use substitutes for the original and harms its market. In Warhol v. Goldsmith, the Supreme Court put substitution at the centre of the first factor too.

That is why the unsealed material is aimed where it is. The defendants’ case is that training is transformative and that chatbots do not replace the news. The publishers now have:

  1. A Microsoft executive calling it theft, several times, in writing.
  2. A Microsoft warning that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”
  3. Measured substitution data from Microsoft itself.
  4. OpenAI’s own head of ChatGPT calling his product “largely substitutive, period.”
  5. Satya Nadella, under deposition earlier in 2026, agreeing that chatbot conversations had substituted going to the underlying source — and testifying that “anything that is paywalled should be licensed by anyone who wants to use it… for grounding or training.” He added that had he known OpenAI scraped paywalled material, he would have invoked Microsoft’s contractual right to require OpenAI to retrain its models.

“The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,” Steven Lieberman, counsel for the New York Daily News, said in a statement after the filing went public.

Two caveats keep this from being a verdict. First, courts have so far leaned toward AI companies on fair use in other cases, and the Trump administration filed a brief in early September 2026 defending OpenAI’s unlicensed use of copyrighted material for training. Second, summary judgment briefing closed on 2 April 2026, and a ruling is expected in the second half of 2026; if any claims survive, trial is projected for 2027. Nothing in the filing establishes infringement — it is a party’s argument, built from the other side’s documents.

But the defendants’ own words are now the strongest exhibit in the plaintiffs’ brief, which is an unusual place for a fair use defence to find itself.

What it could cost, and who is exposed

The numbers explain why both sides are fighting so hard. US statutory damages for willful copyright infringement reach US$150,000 per work. Against hundreds of thousands of unique works, even a modest rate of willfulness findings becomes existential.

There is one court-supervised reference point: Anthropic’s US$1.5 billion settlement with authors in 2025, which works out to roughly US$3,000 per work by the arithmetic widely used since. Applied to the 160,903 unique works in the Project Mango dataset, that is around US$483 million; applied to the 91,692 copies identified in the mid-training sets, roughly US$275 million. Those are back-of-envelope figures, not a judgment, but they show the order of magnitude the discovery fights have been about.

Microsoft and OpenAI are not the only targets. Publishers sued Meta in May 2026 over Llama training data, and dozens of author, artist and newsroom cases now sit in the same MDL. The AI scraping question is no longer a Silicon Valley niche: it is a general liability line item for anyone building a model or a “grounded” AI product on other people’s text.

What AI scraping means for Malaysian publishers and developers

Malaysia is deciding its own version of this question right now, which makes the timing of the unsealed filing unusually relevant.

On 3 July 2026, the Intellectual Property Corporation of Malaysia (MyIPO), an agency under the Ministry of Domestic Trade and Cost of Living, released a public consultation paper on proposed amendments to the Copyright Act 1987. The proposals include a legal framework for the use of copyrighted works in AI training, with emphasis on transparency, fairness and compensation — summarised by MyIPO in terms ordinary creators would recognise: “Creators, publishers and media want to be asked first, paid fairly and told clearly how their works are being used.”

The consultation sets out four broad directions:

  • A purpose-based exception, similar to Japan’s, permitting use for information analysis including AI training, provided it does not unreasonably prejudice the copyright owner.
  • A lawful-access model, similar to Singapore’s, where computational analysis is permitted where the material was lawfully accessed.
  • A hybrid model adapted to Malaysia, often paired with anti-circumvention conditions such as non-excludable contracting and market-harm carve-outs.
  • A full-protection model, requiring express authorisation from the rights holder before AI training.

Written submissions filed with MyIPO through August 2026 split along predictable lines: rights-holder groups warned that a broad text-and-data-mining exception would hand commercial AI developers free rein and create uncertainty for small businesses, while AI and research interests argued against a regime that makes lawful dataset building a litigation risk.

Two Malaysian-specific wrinkles matter for anyone building here:

  1. Copyright is only the first gate. Even a text-and-data-mining exception says nothing about personal data. Training data at scale almost always contains names, faces and user-generated content, and processing that is governed by the Personal Data Protection Act 2010 (Act 709), whose default rule is consent. Malaysian developers need to clear both gates, not one.
  2. Language does not give Malaysian content a pass. AI scraping does not care whether a page is in English, Malay or Chinese. Malay-language sites are still being crawled, indexed and summarised. What differs is that Malaysian publishers have almost no leverage in a licensing negotiation — which is exactly why the MyIPO consultation is the local industry’s best chance to shape rules before the market settles.

There is also a direct commercial echo. If Copilot-style answer engines cut referrals to established publishers by 87% to 93%, the same mechanisms apply to a Malaysian review blog, a school-fees SaaS operator’s content marketing, or a Shopee affiliate guide. Traffic that used to arrive as a click now arrives as nothing.

A practical AI scraping checklist for Malaysian site owners

You cannot litigate this from a hosting plan in Kuala Lumpur. You can, however, make your position explicit and measurable — and that record matters if licensing norms harden the way the MyIPO consultation suggests.

1. Declare your crawler policy in robots.txt. Be explicit about AI training bots rather than relying on a generic rule:

User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: *
Allow: /

Robots.txt is a norm, not a law — it does not stop a determined scraper — but it is the first thing a licensing negotiation or a court looks at when deciding whether access was authorized.

2. Measure who is actually crawling you. Grep your access logs for known AI user agents and get monthly counts, so any future claim of “substantial harm” is backed by your own data:

grep -Ei "GPTBot|ClaudeBot|CCBot|PerplexityBot|Google-Extended" access.log \
  | awk '{print $1, $NF}' | sort | uniq -c | sort -rn | head -20

If you run a WordPress or nginx site, log full user-agent strings for at least 90 days. A shorter window is a common, self-inflicted blind spot.

3. Separate search from training at the server. Blocking a training crawler is different from blocking Googlebot. If you want search visibility but not training use, apply the AI-specific tokens above rather than a blanket block; if your content is the product, invert that and serve a paywall or a licensed feed.

4. Put licence terms in writing. Add explicit AI-training terms to your footer, terms of use and RSS metadata, and keep dated copies. Malaysia’s Copyright Act reforms are heading toward transparency and compensation obligations; being the site that already publishes clear terms puts you on the right side of any “implied licence” argument.

5. Treat your own data pipeline the same way. If you are fine-tuning a model or building retrieval-augmented generation on scraped text, keep a manifest of sources, access method and licence status. The Project Mango and Project Taxi disclosures should be read as a warning about what discovery looks like: dataset provenance becomes the case.

FAQ: The Microsoft ‘largest theft of labor’ filing

Who said AI scraping was “the largest theft of labor in human history”?
Brent Hecht, Microsoft’s Director of Applied Science, in an internal memo, according to the news publishers’ summary judgment brief unsealed on 17 September 2026. He also described the practice as “an astonishing theft of unprecedented proportions.” Microsoft says the views were his own, not the company’s.

Which case is this?
The copyright litigation brought by The New York Times against OpenAI and Microsoft in December 2023, later consolidated with other news plaintiffs — Daily News titles, Ziff Davis, the Center for Investigative Reporting and The Intercept — into MDL No. 25-md-3143 in the Southern District of New York before Judge Sidney H. Stein.

Does the filing prove OpenAI and Microsoft broke the law?
No. It is the plaintiffs’ brief, and some underlying exhibits remain sealed. It does not establish infringement. What it does is put the defendants’ own internal language and Microsoft’s measured click-through data into the fair use argument, where market substitution is the decisive factor.

What is the “doom loop” the filing describes?
An internal Microsoft presentation from January 2024 warned that AI answer engines suppress publisher traffic, which weakens the newsrooms that supply the content, which in turn degrades the model’s own content supply chain. Microsoft’s data showed click-through declines of up to 94% for some publisher groups versus ordinary Bing search.

Is AI scraping legal in Malaysia?
There is no AI-specific exception in the Copyright Act 1987 today. MyIPO’s July 2026 consultation paper proposes four models for AI training, from a Japan-style purpose-based exception to full rights-holder authorisation. Even if a text-and-data-mining exception is adopted, personal data processing under the PDPA 2010 still requires its own lawful basis.

The bottom line

The most damaging evidence in the biggest AI copyright case is not a technical report or an expert model. It is a Microsoft director writing, in a document his own employer produced, that AI scraping of journalism amounts to the largest theft of labor in human history — filed next to Microsoft’s own data showing that its Copilot answer engine cut referrals to those same publishers by as much as 94%.

A summary judgment ruling is expected in the coming months, and a trial in 2027 is possible. Malaysia’s own rulebook is being drafted in parallel. For publishers, developers and site owners here, the practical lesson is the same either way: document what you publish, declare what crawlers may do with it, measure who takes it, and if you build on other people’s data, keep a provenance record you would be comfortable defending in discovery.

More from ZFRBuild