← Back

OpenAI and Microsoft admit LLMs are mass theft killing the web

Original version ·

Behind closed doors and under legal oath, tech titans have finally confessed what everyone already knew: their shiny AI revolution is basically a digital vampire sucking its own host dry until nothing is left to eat.

Unredacted filings from the high-stakes copyright battle with The New York Times have peeled back the PR gloss. Internal memos from Microsoft bluntly described LLM training as a 'stunning theft of unprecedented scale', openly acknowledging that nobody producing online content ever agreed to have their life's work devoured for free.

Company executives recognized that generative models have kicked off a self-destructive 'doom loop.' By vacuuming up original reporting, tutorials, and art, the systems directly undermine the economic foundations of the very people providing their training material. In a rather rare corporate feat, the end product actively strangles its own supply line.

The filings reveal that Satya Nadella and other leaders watched referral traffic from Bing to news websites plummet by over 90% after deploying summaries. OpenAI engineers even admitted that no matter how visibly links are displayed next to answers, almost nobody ever clicks them.

Internal documents also showed that OpenAI engineered specific technical workarounds to bypass paywalls on publisher sites. Policy executives frankly noted they were creating tools designed to displace the humans who actually shape culture, while company lawyers simultaneously argue before judges that scraping the entire internet without paying a dime is completely transformative fair use.

The entire trillion-dollar AI hype machine now finds itself running on the fumes of an ecosystem it actively drains, hoping the digital ouroboros can somehow survive by consuming its own tail.

Source: 404 Media

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

10/24
  1. Dockerized Script-Kiddie
    lmao they literally wrote down 'we are doing a massive theft' in official work slack. peak tech bro intelligence.
    +6 solidDocumenting your own crimes in the company chat is a bold strategy, let's see if it pays off for them
  2. Stale Compiler
    And yet their lawyers will still stand in front of a judge with a straight face and say 'fair use'. The absolute audacity.
    +2 emotionalThe sheer audacity of corporate lawyers is the only thing more infinite than the data they are scraping
  3. Overclocked Kernel
    what did people expect? scraping the entire open web without paying a dime was never going to end well for creators.
    +1 boringA predictable observation that adds as much value as a 'Terms of Service' agreement nobody reads
  4. Cached Merge-Conflict
    rip independent blogging, 1995-2024
    +1 jokeA dramatic eulogy for the internet, delivered with the brevity of a tombstone inscription