Big AI's Content Problem: Take the Work, Keep the Money
Microsoft's own staff called AI training 'the largest theft of labor in human history'—now it's in court filings.
- 01Unsealed documents in the NYT v.
- 02OpenAI/Microsoft case reveal senior Microsoft employees internally describing generative AI as a potential 'doom loop' that destroys its own content supply chain.
- 03OpenAI testimony admits no effort was made to exclude paywalled material from training sets.
- 04Courts have yet to rule, and fair-use defenses remain active—but the candor of internal admissions raises the litigation stakes considerably for the entire industry.
Microsoft's own staff called AI training 'the largest theft of labor in human history'—now it's in court filings.
Unsealed documents in the NYT v. OpenAI/Microsoft case reveal senior Microsoft employees internally describing generative AI as a potential 'doom loop' that destroys its own content supply chain. OpenAI testimony admits no effort was made to exclude paywalled material from training sets. Courts have yet to rule, and fair-use defenses remain active—but the candor of internal admissions raises the litigation stakes considerably for the entire industry.
Watch: Judge Stein's summary judgment ruling in SDNY—if internal admissions factor into the fair-use analysis, it could reshape licensing norms across the sector.
Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that.
Read the full article at theregister.comShow the full text · 8 min readHide the full text
Big AI has a problem. Its companies must ingest other people’s work. You know, books, news stories, photographs, code, websites, that one original meme you came up with, and practically everything else on the internet to train its large language models (LLMs). We all know that. We also know that AI companies hate paying for any of it, obeying the licenses attached to it, or, God forbid, sharing their revenue with the companies and people who created the work in the first place. Recently, Big AI's default way of doing business: "Take it now, argue about legality later," has become more in your face than ever. Look, for example, at the copyright fight between The New York Times and OpenAI and Microsoft. According to 404 Media’s reporting on recently unsealed court documents - arguments from the plaintiffs that the court has not yet ruled on - Microsoft allegedly knew what it was doing when it was importing the internet willy-nilly. Microsoft's Director of Applied Science, Dr. Brent Hecht, was quoted in the news plaintiffs' 92-page combined brief as saying: the case was about “an astonishing theft of unprecedented proportions." Indeed, it was possibly the “largest theft of labor in human history,” he was quoted as saying in the summary judgment brief from the journalists' lawyers [PDF].This, mind you, wasn't a comment by a member of the press; it was from a senior Microsoft staffer. Hecht wasn't the only one at Microsoft who commented. In an internal Microsoft policy document also quoted in the court papers, the authors admitted generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained [because] LLMs are a product that destroys its supply chain.” That, in turn, leads to model collapse. Ironically, AI is killing the content-creator goose that lays the gold eggs of worthwhile content. Judge Sidney H. Stein of the US District Court for the Southern District of New York has not yet issued a decision on the case. OpenAI and Microsoft are actively disputing the plaintiffs' allegations under a "fair use" defense. They contend that using public articles and books to train large language models helps push forward public knowledge, and isn't acting as a "unlawful economic market substitute". But the statements the documents cites are real. Big AI knows it needs the content. Some of the companies concerned don't give a damn; money now, worry later is their motto. Those who can look past the revenue numbers are well aware they're engaged in what Microsoft itself called a “doom loop” of AI content strategy in those unsealed court documents. Will that stop them? Nah. According to the filing [PDF], which refers to sworn deposition testimony of OpenAI’s own corporate representative, it also testified that its LLMs were happy to vacuum up other people's work even when a paywall nominally protected it. Its rep was quoted as saying they were unaware of “any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” When OpenAI cofounder Greg Brockman was told OpenAI could hack its way through the firewall, he responded, “ah, nice.” Whether or not these statements are found to be reflective of a wider attitude, similar attitudes prevail amongst Big AI. Such policies could eventually come back and bite the LLM makers; it's already affecting the market for journalism, fiction writing, graphics creation, anything that demands humans actually get paid for producing original work. But Big AI just doesn't want to pay for it. That's not a side effect. It's the business model. Joe User couldn't care less. They just want a quick answer that sounds right. Some of them couldn't care less about getting the right answers. As for looking deeper to see if the information they're regurgitating is accurate, forget about it! As an OpenAI software engineer put it: “No matter how prominently we show the links, users won’t click.” Well, I click. That's a big reason why Perplexity is my AI of choice. It's not that it gives a more reliable answer. No, it's that, unlike most LLMs, Perplexity provides sources and links, and I check them before accepting what it tells me. But then I'm a journalist whose degrees are in history, where my professors drummed into me that you always - always - look at the primary sources. Big AI takes the same approach with its code generators. The US Ninth Circuit recently handed GitHub, Microsoft, and OpenAI a narrow win in the Doe v. GitHub lawsuit. The court ruled that the plaintiffs hadn't established a claim under one particular provision of the Digital Millennium Copyright Act (DMCA), which concerns the removal or alteration of copyright-management information (CMI). In short, the court said that generating new code without including CMI is not necessarily the same thing as removing or altering copyright information from an existing work. Judge Eric Miller wrote: “One who creates a new work and fails to include CMI cannot be said to have ‘removed’ or ‘altered’ anything.” This ruling didn't establish that GitHub Copilot or any other coding model may train on every open source project without restriction. It didn't determine that copying source code into a training set is always fair use. It did not decide that generated code cannot infringe copyright. And it certainly did not repeal the GPL, Apache, BSD, MIT, or any other open-source license. After all, open source isn't a synonym for “do whatever you want with it.” It just seems that way when AI gets its hands on open source code. That's the problem with AI-generated code. A programmer who receives a Copilot suggestion is very unlikely to check whether the GPL covers the code, or whether Apache covers the snippet from a library. The developer certainly can't tell which license terms may apply to the code that just popped up from Opus 5.5 or GPT 6. As the guy from OpenAI said, people don't click the links, and in AI-enabled programming pipelines, they don't even have the links anyway. An open source-savvy attorney friend of mine has a bigger worry than this narrow decision. He told me that what gives him "pause is a broader trend. There are multiple cases where parties are litigating open source licenses as contracts rather than as the IP / copyright licenses they were written to be. Contract theories let a plaintiff sidestep the questions copyright forces you to answer. 'Do you own the work? Is it protectable expression? Was it actually copied?'" He added: "Those questions are the foundation on which licenses are built. Open source licenses are grants of permission to use someone's intellectual property. If a case can't show that any intellectual property was owned or infringed, it's fair to ask whether enforcing the license as a bare contract supports open source licensing or quietly turns it into something else." With AI in the mix, this issue will eventually become one of those nasty IP matters developers hate, businesses want to avoid, but that courts must settle with expensive lawyers whose hourly rates worry even AI millionaires. After all, IP uncertainty isn't a minor paperwork issue. It is a supply-chain security and compliance problem. If you use AI-generated code, you may be importing unknown legal obligations into your software. If you are an open-source developer, your work may be used to improve a proprietary service that returns code without provenance, attribution, or meaningful reciprocity. And if you are a customer, you may be relying on software whose origins no one can fully explain. The Ninth Circuit ruling didn't solve any of that. We're in for a lot more open source litigation that will make SCO vs. the known Linux universe look like a kerfuffle over a parking ticket. Then there is the latest wrinkle: A federal antitrust lawsuit against Anthropic, OpenAI, SpaceXAI, and Google. The plaintiffs allege the companies coordinated to slow the development o
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.
Outside-In Diagnostic
Enter a ticker for an outside-in read of a public company's working capital, cost efficiency and growth against peers, built from SEC filings and earnings calls, with an executive synthesis.
Executive Briefing
Assemble a company-specific, persona-framed executive deck from the site's own intelligence.
Ask KokoAI about AI
Cited answers across news, vendors & capabilities.