Photo by Steve A Johnson via Pexels

OpenAI and Microsoft Knew AI Would Break the Web

4 Min Read

Newly unsealed court documents from a major lawsuit against OpenAI and Microsoft have pulled back the curtain on something many suspected but few could prove: both companies understood, in their own words, that building large language models on scraped web content would create a catastrophic feedback loop for the broader internet. They pressed forward anyway.

The internal records are striking in their candor. One Microsoft document warned that the company’s AI content strategy had started a doom loop that would simultaneously hurt model performance and damage the entire web. The document noted that it is highly unusual for a product to threaten the economic foundations of its own essential suppliers, yet that is precisely the situation these companies created. That is not a critic talking. That is an internal memo.

The Scale of What Was Taken and What Was Lost

A Microsoft director of applied science described the harvesting of data to train AI models as the largest theft of labor in human history, and said that the company’s legal defense made a complete mockery of fair use principles. Meanwhile, internal OpenAI communications acknowledged that GPT-4 had memorized enormous amounts of data and would be, in their own phrasing, insanely good at regurgitation. Several examples in the court filing show the AI reproducing long passages verbatim from copyrighted articles.

The downstream consequences are not hypothetical. OpenAI’s own media and economic experts estimated that AI-generated summaries have contributed to referral traffic declines of as much as 60 percent for major publishers. For an industry already operating on thin margins, that kind of drop is existential. The web’s economic model has long depended on clicks, and when users get complete answers from a chatbot, one OpenAI executive admitted there is simply no good reason to click through to the original source.

A Supply Chain That Eats Itself

Perhaps the most striking admission is that Microsoft internally acknowledged LLMs are a product that destroys its own supply chain. If AI tools replace the need to visit news sites, blogs, and specialty publishers, those outlets will produce less content. Less original content means lower quality training data for future models. The very tools designed to synthesize human knowledge are undermining the conditions that make new knowledge creation economically viable.

This is not a bug that escaped detection. It was flagged internally and documented. The pursuit of what one executive described as gazillions of dollars in commercial potential was prioritized over a sustainable content ecosystem.

What This Means for People Buying AI-Powered Products

For consumers evaluating AI tools, subscriptions, and enterprise software right now, this context matters. Products built on these foundations carry real reputational and regulatory risk as litigation expands globally. Buyers investing in AI-driven productivity platforms should ask vendors directly how their models are licensed and whether training data practices are legally defensible. The market is shifting fast, and companies that built responsibly from the start are increasingly becoming the smarter long-term investment.

Share This Article