Microsoft Called AI Scraping 'Theft.' It Scraped Anyway.

Microsoft Called AI Scraping 'Theft.' It Scraped Anyway.

Unsealed filings show Microsoft privately labeled OpenAI's data practices theft while both companies scraped paywalled Times content and warned internally it would gut publishers. The evidence shifts the legal and commercial ground under every AI lab that trained on licensed journalism.

Microsoft's own executives called AI data scraping 'the largest theft of labor in human history' β€” then Microsoft and OpenAI kept scraping paywalled New York Times journalism anyway. Newly unredacted court filings turn that contradiction into evidence, not commentary.
  • Newly unsealed filings show Microsoft executives privately described AI data scraping as 'the largest theft of labor in human history' while Microsoft and OpenAI continued scraping paywalled New York Times content.
  • The documents include internal warnings that the practice would gut publishers β€” evidence of intent, not just impact.
  • The key tension: Microsoft is simultaneously OpenAI's largest investor and, in these filings, a critic of OpenAI's data practices.
  • This article explains what the filings change legally and commercially, and what publishers, labs, and enterprises should do next.

What exactly did the unredacted filings reveal?

According to TechCrunch, the newly unsealed court filings show Microsoft privately called OpenAI's data practices 'theft' while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers. The headline quote β€” 'the largest theft of labor in human history' β€” was written by a Microsoft executive, per the filings as reported by TechCrunch. The significance is not the sentiment. Everyone in media suspected the labs knew. The significance is documentation. Internal warnings that scraping would gut publishers establish that the companies understood the harm and proceeded anyway. In litigation, that is the difference between negligence and knowledge β€” and knowledge is what supports claims for willful infringement and, potentially, statutory damages. TechCrunch reported the filings on September 17, 2026. The underlying case is The New York Times' copyright suit against Microsoft and OpenAI. The unsealing matters because redactions previously hid the most quotable internal material.

Who is actually exposed β€” Microsoft, OpenAI, or both?

Both, but asymmetrically. OpenAI built the models; Microsoft built the pipeline and, per the filings, the internal critique. Microsoft's dual position is the most awkward fact in the record: it is OpenAI's largest investor, its cloud provider, and β€” according to these documents β€” an internal critic of OpenAI's data sourcing. For OpenAI, the exposure is direct: the Times alleges its paywalled content was scraped and used in training. For Microsoft, the exposure runs through both its investment and its own scraping activity described in the filings.
Microsoft Called AI Scraping Theft. It Scraped Anyway.

What are the operational tradeoffs for AI labs that trained on scraped data?

The tradeoff is now explicit: cheap data versus clean provenance. Scraping paywalled journalism cost the labs almost nothing and produced models fast. Licensing that same corpus is expensive, slow, and requires per-publisher negotiation. The practical consequence is a bifurcated market. Frontier labs with capital will sign licensing deals and pass costs to enterprise customers. Smaller labs without licensing budgets will either train on lower-quality open corpora or accept legal risk. Neither outcome is free. TechCrunch's reporting makes clear the internal Microsoft warnings anticipated exactly this publisher-gutting dynamic. The filings do not show the companies stopping. They show them continuing while aware.

What should publishers and enterprise AI buyers do next?

Publishers should treat these filings as negotiating leverage, not just evidence. The documented internal warnings raise the labs' litigation risk, which raises the value of a settlement that includes licensing terms. The New York Times is already positioned to extract terms; smaller publishers should consider collective licensing vehicles rather than individual deals. Enterprise AI buyers should start asking vendors for data provenance documentation. If a model was trained on disputed corpora, the buyer inherits reputational and, in some jurisdictions, legal exposure. Procurement teams that do not ask will find out during discovery.
PartyPosition in the filingsImmediate riskBest move
MicrosoftInvestor, cloud provider, internal critic of OpenAI scrapingWillful-infringement exposure; governance credibilitySettle with licensing terms; separate its data pipeline from OpenAI's
OpenAIModel builder; alleged scraper of Times contentStatutory damages; injunction risk on training dataSign publisher licenses; document provenance going forward
The New York TimesPlaintiff with documented internal warnings as evidenceNone material; upside in settlement valuePush for licensing terms plus damages, not just cash
Other publishersNon-parties with similar claimsBeing left out of any settlement termsForm collective licensing coalitions now
Enterprise AI buyersDownstream users of disputed modelsReputational and contractual exposureRequire provenance documentation in vendor contracts
VerdictThe New York Times and publisher coalitions hold the strongest hand; Microsoft's dual role makes it the most exposed defendant.

The unsealed filings prove that AI labs knowingly built commercial products on stolen paywalled journalism, and the only durable fix is court-enforced licensing β€” not voluntary pledges.

Short term, this is a settlement-value event. The Times now has internal documents showing Microsoft executives understood the harm. That raises the price of resolution and makes a pure cash settlement less likely; licensing terms become the currency. Long term, this accelerates the shift from scraping to contracting. Labs that built models on disputed corpora will spend the next several years retroactively licensing or rebuilding datasets, and the cost curve for frontier training rises accordingly.

The winners are publishers with archives worth licensing and the lawyers who structure those deals. The losers are smaller AI labs that cannot afford licensing and cannot afford litigation. Microsoft loses the most in credibility: it cannot simultaneously argue it was OpenAI's responsible partner and that OpenAI's data practices were theft.

Concrete prediction: by mid-2027, Microsoft will announce a paid content-licensing framework covering multiple major publishers, structured to reduce its willful-infringement exposure in the Times case. Watch for the first deal within two quarters of the next substantive court ruling.

What are the falsifiable predictions?

1. Microsoft will announce a multi-publisher content-licensing framework by mid-2027, explicitly covering news archives, as part of its litigation posture in the Times case. 2. At least one additional major publisher will file a copyright suit against OpenAI or Microsoft citing the unsealed filings as evidence of knowledge, within 12 months. 3. Enterprise AI procurement contracts will begin including data-provenance warranty clauses by 2027, driven by buyer risk teams rather than regulators.
  1. December 2023
    The New York Times sues Microsoft and OpenAI

    The Times files a copyright infringement suit over the use of its paywalled journalism in AI training.

  2. 2024-2025
    Discovery and redactions

    Internal Microsoft and OpenAI documents are produced in discovery, with key passages redacted from public view.

  3. September 2026
    Filings unsealed

    Newly unredacted filings reveal Microsoft executives privately called OpenAI's data practices 'theft' and warned scraping would gut publishers.

Estimated annual content-licensing spend by major AI labs (estimated)

What should readers remember after closing this tab?

  • The quote is not the story; the documentation of intent is. Internal warnings that scraping would gut publishers are what raise legal exposure.
  • Microsoft's dual role β€” investor and internal critic β€” is its biggest liability and its weakest defense.
  • The operational shift is from scraping to licensing, which raises the cost of frontier training and favors well-capitalized labs.
  • Publishers should negotiate collectively; enterprise buyers should demand provenance documentation now.
  • The Times holds the strongest hand, and licensing terms β€” not just cash β€” will define any settlement.
Microsoft exec called AI scraping β€˜the largest theft of labor in human history,’ new unredacted filings reveal
Embedded source image Source: techcrunch.com. Original reporting.

Source and attribution

TechCrunch AI
Microsoft exec called AI scraping β€˜the largest theft of labor in human history,’ new unredacted filings reveal

Discussion

Add a comment

0/5000
Loading comments...