OpenAI, Microsoft knew AI could hurt news publishers in major way, unsealed US court filings reveal
RNA Media illustration for representation.
New Delhi: Newly unsealed United States court filings have revealed that senior personnel at Microsoft and OpenAI privately recognized the potential harm that large-scale use of news content for artificial intelligence training could cause to publishers, while the companies continued to defend the practice as lawful under the US doctrine of fair use. The disclosures have emerged in the copyright litigation brought by the New York Times and other publishers, adding significant detail to a legal dispute that has become a test of how copyright law applies to generative AI.
The filings cited by the publishers indicate that OpenAI used millions of news articles in datasets assembled for training and developing its AI systems, with NYT alleging that more than 10 million articles were scraped and that nearly one-third came from its own website. The material also refers to internal Microsoft discussions in which Brent Hecht, the company’s director of applied science, described the scale of the copying in exceptionally stark terms and warned that the technology could damage the wider information ecosystem on which AI systems themselves depend.
According to the unsealed material, Hecht warned in an internal document that the AI industry’s dependence on large quantities of online content could create a damaging feedback loop, because AI-generated answers could reduce traffic to the very publishers whose work supplies the underlying information. Microsoft data cited in the litigation reportedly showed that its Copilot service could produce a substantial fall in click-throughs to The New York Times compared with conventional Bing search, an issue the publishers argue is central to their claim that AI products can compete directly with original journalism.
The documents also contain comments attributed to OpenAI personnel about the threat posed to publishers as chatbots became increasingly capable of delivering information directly to users. OpenAI executives and employees are alleged to have recognized that products such as ChatGPT could substitute for visits to news websites, an argument the publishers are now using to challenge the companies’ contention that training AI models on copyrighted works is sufficiently transformative to qualify for fair-use protection.
Another contentious aspect concerns the way some copyrighted material was allegedly obtained. The unsealed filings describe efforts by OpenAI researchers to access material behind NYT’s paywall and allege that large collections of news content were assembled from web indexes and datasets, while other evidence cited by the publishers points to the removal of copyright notices from some training material. These allegations are being disputed in the broader litigation and should not be treated as established findings of fact unless and until the court rules on them.
The scale of the material described in the filings is considerable, with the publishers alleging that datasets used in the development of AI systems contained tens of thousands of copies of works from news organizations. One dataset assembled through a project referred to as Project Mango is alleged to have contained more than 160,000 unique works from the publishers involved in the litigation, while another Common Crawl-derived collection reportedly included more than two million documents from nytimes.com alone.
Microsoft, however, has rejected the suggestion that comments made by individual employees establish the company’s legal position. A Microsoft spokesman said the internal documents cited by NYT were authored by Hecht in his capacity as a research academic and did not represent Microsoft’s position, which remains that its use of copyrighted material for AI development can constitute transformative use protected by copyright law. OpenAI did not respond to some requests for comment reported by news organisations.
The legal question is particularly important because US copyright law does not contain a specific rule that automatically permits or prohibits the use of copyrighted works to train generative AI models. The companies argue that training involves transforming existing material into models that perform a different function, while publishers contend that unlicensed copying is unlawful and that AI systems can ultimately substitute for the original works and undermine the market for journalism.
The dispute has also acquired a wider policy dimension. Earlier this month, the US justice department filed a brief supporting the position that OpenAI’s use of NYT’s articles for AI development did not violate copyright law, invoking considerations including technological progress, economic growth and national security.
For the news industry, the case goes beyond the question of compensation for individual articles. At its core is a more consequential issue: whether companies developing foundation AI models can commercially ingest vast quantities of professionally produced journalism without obtaining licences, particularly when the resulting systems may answer questions that previously sent readers directly to publishers’ websites.
The court’s eventual ruling could therefore have implications well beyond the parties to this lawsuit, potentially influencing licensing arrangements, AI training practices and the economic relationship between technology companies and publishers. The plaintiffs have sought summary judgment, but a final determination on the competing copyright arguments remains pending, with a ruling not expected until 2027.
