New unredacted information in a copyright lawsuit filed by The New York Times against OpenAI and Microsoft three years ago reveals an admission that AI scraping is tantamount to theft and that AI products pose a serious threat to publications.
According to the complaint, Microsoft executives privately described the company’s AI training practices as “theft,” and OpenAI’s own executives said their AI models posed an “existential threat” to the publishers and journalists who were trained on the job.
The unsealed documents also detail how these companies allegedly obtained and used that content by circumventing paywalls undetected, building training datasets through mass scraping, and intentionally removing copyright notices from training data.
It’s worth noting that much of the new information comes from the Times’ own summaries rather than the underlying exhibits, which remain sealed. The following quotes are presented out of their original context.
The unredacted filing is the latest escalation in a three-year-old lawsuit in which The New York Times initially alleged that the companies violated copyright law by training generative AI models based on the content.
There is no clear answer to the question of whether AI companies can legally use copyrighted material to train AI, but judges have generally favored AI companies’ arguments that the training constitutes “fair use.” This legal rule allows the unauthorized use of copyrighted works in certain cases, such as parody, news reporting, and criticism. Earlier this month, the Trump administration contributed a brief defending OpenAI’s unauthorized use of copyrighted material to train its LLMs.
However, some of the new approvals violate OpenAI’s fair use defenses, specifically the rule’s requirement that the use do not substitute for or harm the market for the original work.
For example, Microsoft’s own data shows that Copilot’s “response engine” reduced click-through rates for the New York Times domain by 93% compared to traditional Bing search. An internal Microsoft presentation written by Brent Hecht, Microsoft’s director of applied science, in January 2024 described this decline as a “loop of doom” that “simultaneously negatively impacts our models and the performance of the web as a whole.”
“While it is highly unusual for an end product to threaten the economic foundations of its critical suppliers, we created this situation for our LLM business with respect to its ‘content supply chain,’” the Microsoft document cited in the filing reads.
Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anyone who wants to use something behind a paywall for grounding and training purposes should license it,” and declared that if he had “known that OpenAI was collecting and training on information behind a paywall,” he “would have invoked Microsoft’s right to require OpenAI to retrain its models.”
Some of the content violated various pillars of the fair use test. Nick Turley, head of ChatGPT at OpenAI, wrote in an internal communication that publishers face an “existential threat” from products like chatbots, which are “nearly substitutable” and “will become even more substitutable as they improve.”
OpenAI President Greg Brockman called the model “great for news.” Earlier this year, Nadella agreed to a declaration that conversations with chatbots “replace the provision of information on a website on an AI platform without the need to access the underlying information source.”
This kind of language speaks to how technology can directly compete with the original work rather than transform it.
Microsoft’s document says there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data based on the training of the underlying model.”
The scale of the copy is astonishing. This document reveals for the first time that OpenAI’s intermediate training dataset alone contains over 91,692 copies of copyrighted works published by the NYT, the Daily News, and the Center for Investigative Reporting. The dataset derived from Common Crawl contained over 2 million documents from nytimes.com alone.
In a January 2023 internal memo, Hecht called it a “stunning theft of unprecedented scale” and “the largest theft of labor in human history.”
The filing provides new details about how OpenAI and Microsoft obtained the plaintiff’s content, including scraping it from Bing Index.
“OpenAI provided the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its commercial products,” the filing states. “Microsoft has similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”
The companies allegedly assembled Project Mango data into a training dataset that included at least 160,903 copies of original works from news publishers.
To get the most out of scraping, OpenAI employees allegedly devised a plan to bypass paywalls without being detected. According to the filing, OpenAI researcher Nick Rider told Brockman about “hacks to get around paywalls at any time,” to which Brockman responded, “Oh, that’s nice.”
OpenAI employees also allegedly built training datasets such as WebText and WebText2 that relied heavily on scraped news content. It is also said to have pulled millions of articles from Common Crawl, a free and open repository of web crawl data. “The findings also explain a deliberate effort to remove copyright notices from the training data before reaching the model, because we don’t want the model to output a ‘copyright notice’ to the user,” the researchers wrote.
OpenAI and Microsoft did not respond to requests for comment.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
