Microsoft and OpenAI Executives Called AI Training ‘Astonishing Theft’ in Internal Documents

Internal Admissions Expose AI Industry’s Stance on Content Scraping

Newly unredacted court documents have exposed damning internal communications from Microsoft and OpenAI executives who privately acknowledged that their AI training practices constituted theft on an unprecedented scale. The revelations emerged from the copyright lawsuit The New York Times filed against both companies three years ago, with the unsealed material containing admissions that directly contradict the firms’ public legal defenses.

Brent Hecht, Microsoft’s Director of Applied Science, repeatedly warned in internal documents that scraping news content for AI training represented “an astonishing theft of unprecedented proportions,” according to the plaintiff’s motion. Hecht went further, calling the practice perhaps the “largest theft of labor in human history” in what news organizations describe as explosive testimony that undermines the tech giants’ fair use arguments.

The unsealed filing represents the latest escalation in the three-year legal battle, in which The New York Times initially alleged both firms violated copyright law by training generative AI models on its content without permission. The question of whether AI companies can legally use copyrighted material to train their models remains unresolved, though judges have largely favored AI firms’ arguments that such training constitutes fair use-a legal doctrine permitting certain unauthorized uses of copyrighted work for purposes like parody, news reporting, or criticism.

Executive Warnings About Existential Threats to Publishers

Beyond the theft admissions, OpenAI’s own leadership acknowledged the existential threat its AI models posed to the publishers and journalists whose work trained them. Nick Turley, head of ChatGPT, wrote in an internal message that publishers would face an “existential threat” from commercial products trained on news content that could substitute for news providers. OpenAI cofounder and president Greg Brockman reinforced this assessment, stating that generative AI products “are largely substitutive, period, [and] will get more and more substitutive as they get better.”

These internal acknowledgments prove particularly significant because several of the newly revealed admissions run counter to OpenAI’s fair use defense, especially the legal rule’s requirement that use doesn’t substitute for or harm the market for the original work. The plaintiff’s counsel argues these admissions “eviscerate” any claim to fair use-a position with which Hecht apparently agrees, as he recognized that prevailing on this defense would “make a complete mockery of the idea of ‘fair use.'”

Documented Evidence of Market Harm and the ‘Doom Loop’

Microsoft’s own data demonstrates concrete market harm to news publishers, with the company’s Copilot “answer engine” causing click-through rates for The New York Times’ domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Hecht in January 2024 describes this decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,'” the Microsoft document states, as quoted in the filing.

The unsealed material also details how both companies allegedly obtained and used publisher content through systematic evasion tactics, including bypassing paywalls undetected, building training datasets via mass scraping, and deliberately stripping copyright notices from training data. Much of the new information comes from The Times’ own brief rather than the underlying exhibits, which remain sealed, and the quotes are presented without their original context.

Legal Landscape and Timeline

The damning documents emerge during the middle phase of a case that began in 2023, when The New York Times accused both AI labs not only of vacuuming up its original news content, but of using that content to train AI products meant to compete directly with traditional news outlets. The plaintiff’s accusations allege the firms violated copyright law by copying “millions” of copyrighted articles in their entirety without permission to produce substitutive commercial AI products.

Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its large language models, adding a political dimension to the ongoing legal debate. The presiding judge is expected to decide whether to allow the case to proceed to trial by sometime in 2027, according to recent reports.

Broader Implications for the AI Industry

The threat of perjury has forced AI executives to articulate concerns they kept private during the industry’s rapid expansion. The internal acknowledgments by Microsoft and OpenAI leadership suggest that top officials understood the legal and ethical implications of their training methods even as they publicly defended those practices. Hecht’s characterization of the AI buildout as the “largest theft of labor in human history” represents an extraordinary admission that attorneys for news plaintiffs are using to argue their case.

For years, both Microsoft and OpenAI fought to keep certain information out of the public eye in their battle with news organizations that accused the AI firms of teaming up to violate copyright laws by stealing vast quantities of news content to train AI systems. The details that should never have been marked confidential are now starting to emerge, providing news groups with evidence they allege shows exactly how both companies viewed the threat to journalism before unleashing new AI products like ChatGPT and Copilot.

The battle remains far from over, with significant legal questions about AI training practices, copyright law, and fair use doctrine still unresolved. However, the newly unsealed admissions provide news plaintiffs with powerful ammunition in their argument that the AI industry built its foundation on what one executive called an unfathomably large pile of stolen intellectual property.