‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft

‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft

Executives working on AI at Microsoft and OpenAI admitted what its critics have been saying all along: Large language models are predatory pieces of technology that have been built on what a Microsoft executive called “an astonishing theft of unprecedented proportions,” and the “largest theft of labor in human history.” An internal Microsoft document said generative AI products have created a “doom loop” that is killing “the entire web.” 

Those statements and a series of other mask-off moments feature heavily in an unredacted court filing that was unsealed Thursday in the behemoth New York Times vs OpenAI copyright lawsuit that has been winding its way through the court system for years. In a filing asking for summary judgment (basically, a filing with the court asking it to rule), lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained either sealed or redacted at the request of Microsoft and OpenAI. 

‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft

It’s easy to see why the AI companies wanted to hide this from the public. The statements, taken together, are some of the most damning indictments of the ways LLMs were trained, how they worked, and the immediate threat they pose to human labor. It is a reminder that even as AI becomes more powerful and companies try to shift the narrative to the supposed existential risk of “superintelligent” AI, the tools they have already built were created by stealing from human creativity and labor and are by definition existential threats to the human labor market.

“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” an internal Microsoft document cited in the case read, adding “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

The filing was written by lawyers for the New York Times but is largely comprised of statements and interviews with big tech executives that admit both that LLMs are largely trained on stolen content, that they represent an existential risk for the human writers, artists, and media companies that they stole from, and that their products have started a “doom loop” that is eating the web and destroying the businesses that these companies stole from. 

The unredacted filing was found by Jason Kint, the CEO of Digital Content Next, a trade organization that represents digital media companies. Kint has been closely following and posting about massive AI copyright lawsuits.

The court filing cites an internal Microsoft document that found that AI products steal from human content creators, then cannibalize clicks from the people and websites they’ve stolen from, thereby destroying their business models. 

“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” the document said. 

Microsoft executives, including CEO Satya Nadella, testified under oath that after ripping content from the New York Times and other news sites, clicks to those news sites fully cratered, falling by more than 90 percent on Bing. 

Documents obtained during the court proceedings found that OpenAI created “a hack to get around nytimes paywall,” to which OpenAI cofounder Greg Brockman said “ah, nice.” Microsoft executive Brent Hecht wrote that LLMs steal content “without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content.” 

OpenAI’s policy director Jack Clark wrote that the company was “creating systems that substitute for the labor of the people that define the ‘culture’ of society” and Microsoft, in a policy document, wrote that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained […] LLMs are a product that destroys its supply chain.” OpenAI called itself an “existential threat” to news publishers, and an OpenAI software engineer testified that “no matter how prominently we show the links, users won’t click.” 

None of this is at all surprising to anyone who has been paying attention to the development of generative artificial intelligence, but the document, taken in whole, is a real they-admit-it situation. OpenAI’s and Microsoft’s lawyers have been trying to argue that their model training is fair use and transformative under copyright law and that they are building something that is fundamentally different from the human labor that it was trained on. But internally, these executives know that what they have built has been built on stolen content and that the products they’ve made are cannibalizing the sources they’ve stolen from and destroying the internet as we know it.   

Scroll to Top