Trending: On-device modelsSearch
iHeartGeek
iTECH

Unsealed filings show Microsoft and OpenAI knew what AI was doing to news

A less-redacted court filing in the OpenAI copyright case quotes Microsoft and OpenAI staff warning that AI scraping was 'the largest theft of labor in human history'.

A thick stack of newspapers on a desk with glowing blue strands of light threading through the pages, beside a blurred computer monitor and keyboard

The long-running copyright fight between news publishers and the companies that build large language models has been conducted mostly in redactions. On 17 September, that changed a little. The news plaintiffs suing OpenAI and Microsoft filed a public, less-redacted version of their summary-judgment brief in the Southern District of New York, and the passages they unsealed are the companies' own staff talking about what AI meant for journalism.

What the unsealed filing says

The brief runs to 92 pages and it is blunt. It quotes Microsoft's Director of Applied Science calling the copying of news archives 'an astonishing theft of unprecedented proportions', and says the same record describes it as perhaps the 'largest theft of labor in human history'. OpenAI's Head of ChatGPT wrote that publishers face an 'existential threat' from AI products that, in the plaintiffs' quotation, 'are largely substitutive, period' and 'will get more and more substitutive as they get better'. The plaintiffs' argument is that admissions like those strip the fair-use defence of its foundation, because fair use turns on whether a copy substitutes for the original.

Microsoft's own document called it a 'doom loop'

The most quotable passage is a Microsoft document the plaintiffs reproduce: 'Our AI content strategy has started a doom loop that will hurt the performance of our models and the entire web at the same time,' it reads, adding that 'it is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain.' The filing backs that up with Microsoft's own numbers: 83 to 93 per cent drops in click-through rates for The Times' and the Daily News plaintiffs' domains, and 51 to 94 per cent for the Ziff Davis titles, when readers met Copilot's answer engine instead of ordinary Bing results. An OpenAI engineer, quoted in the same brief, put the same idea more simply: 'no matter how prominently we show the links, users won't click.'

Paywall 'hacks', datasets and an abandoned fix

The memorandum also revisits how the content was gathered. It quotes an internal exchange in which Nick Ryder told OpenAI co-founder and president Greg Brockman about 'a hack to get around nytimes paywall', and records Brockman's two-word reply: 'ah nice'. It cites a training dataset whose top domains included 2,182,079 articles from chicagotribune.com and 2,064,805 from nytimes.com, and it notes that OpenAI announced a media manager tool in 2024 to let publishers express how their work should be used, then abandoned the project. Microsoft chief executive Satya Nadella, the filing says, testified that anything behind a paywall should be licensed, and that he would have made OpenAI retrain its models had he known paywalled material had been scraped.

What the other side says

OpenAI and Microsoft filed their own summary-judgment motions on 4 September and both deny infringement and rely on fair use. Their experts argue that individual news sources are 'entirely fungible' or 'largely interchangeable' for training, which is a technical claim as much as a legal one. Nothing has been decided: the memoranda are competing versions of the same record, and the quotations the plaintiffs have just unsealed are their characterisation of documents the court has not yet weighed. Separately, Bloomberg Industry Group asked the court on 17 September for permission to intervene for the limited purpose of unsealing and un-redacting more of the file, which suggests the public version of this litigation is still only partly visible.

Our opinion

The striking thing about these passages is not that a Microsoft engineer privately understood what scraping does to the businesses that supply the raw material. It is that the understanding was written down, circulated and filed while the same companies were telling regulators and publishers that AI would send traffic back. A doom loop is a systems problem, and the document does not blame a rogue crawler; it blames a strategy. That is why the fair-use argument now looks shakier than the technology: you cannot claim a transformative public benefit when your own file records the harm and prices the fix. Publishers should still be careful what they wish for, because a licensing regime run by three or four model owners would concentrate British and American journalism under the same handful of balance sheets it is complaining about. But the choice is no longer whether news gets paid for, only who does the paying.