Close Menu
  • Home
  • AI
  • Art & Style
  • Economy
  • Entertainment
  • International
  • Market
  • Opinion
  • Politics
  • Sports
  • Trump
  • US
  • World
What's Hot

Even the British King is hesitant about AI.

September 17, 2026

How will the Fed’s interest rate hikes ripple through global markets?

September 17, 2026

What an Oscar-winning movie can teach us about investing through the AI slowdown debate

September 17, 2026
Facebook X (Twitter) Instagram
Smart Breaking News on AI, Business, Politics & Global Trends | WhistleBuzz
Facebook X (Twitter) Instagram
  • Home
  • AI
  • Art & Style
  • Economy
  • Entertainment
  • International
  • Market
  • Opinion
  • Politics
  • Sports
  • Trump
  • US
  • World
Smart Breaking News on AI, Business, Politics & Global Trends | WhistleBuzz
Home » New unedited filing reveals Microsoft executive calls AI scraping ‘the greatest labor theft in human history’
AI

New unedited filing reveals Microsoft executive calls AI scraping ‘the greatest labor theft in human history’

Editor-In-ChiefBy Editor-In-ChiefSeptember 17, 2026No Comments5 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email Copy Link
Follow Us
Google News Flipboard
Share
Facebook Twitter LinkedIn Pinterest Email


New unredacted information in a copyright lawsuit filed by The New York Times against OpenAI and Microsoft three years ago reveals an admission that AI scraping is tantamount to theft and that AI products pose a serious threat to publications.

According to the complaint, Microsoft executives privately described the company’s AI training practices as “theft,” and OpenAI’s own executives said their AI models posed an “existential threat” to the publishers and journalists who were trained on the job.

The unsealed documents also detail how these companies allegedly obtained and used that content by circumventing paywalls undetected, building training datasets through mass scraping, and intentionally removing copyright notices from training data.

It’s worth noting that much of the new information comes from the Times’ own summaries rather than the underlying exhibits, which remain sealed. The following quotes are presented out of their original context.

The unredacted filing is the latest escalation in a three-year-old lawsuit in which The New York Times initially alleged that the companies violated copyright law by training generative AI models based on the content.

There is no clear answer to the question of whether AI companies can legally use copyrighted material to train AI, but judges have generally favored AI companies’ arguments that the training constitutes “fair use.” This legal rule allows the unauthorized use of copyrighted works in certain cases, such as parody, news reporting, and criticism. Earlier this month, the Trump administration contributed a brief defending OpenAI’s unauthorized use of copyrighted material to train its LLMs.

However, some of the new approvals violate OpenAI’s fair use defenses, specifically the rule’s requirement that the use do not substitute for or harm the market for the original work.

For example, Microsoft’s own data shows that Copilot’s “response engine” reduced click-through rates for the New York Times domain by 93% compared to traditional Bing search. An internal Microsoft presentation written by Brent Hecht, Microsoft’s director of applied science, in January 2024 described this decline as a “loop of doom” that “simultaneously negatively impacts our models and the performance of the web as a whole.”

“While it is highly unusual for an end product to threaten the economic foundations of its critical suppliers, we created this situation for our LLM business with respect to its ‘content supply chain,’” the Microsoft document cited in the filing reads.

Microsoft CEO Satya Nadella also testified in a deposition earlier this year that “anyone who wants to use something behind a paywall for grounding and training purposes should license it,” and declared that if he had “known that OpenAI was collecting and training on information behind a paywall,” he “would have invoked Microsoft’s right to require OpenAI to retrain its models.”

Some of the content violated various pillars of the fair use test. Nick Turley, head of ChatGPT at OpenAI, wrote in an internal communication that publishers face an “existential threat” from products like chatbots, which are “nearly substitutable” and “will become even more substitutable as they improve.”

OpenAI President Greg Brockman called the model “great for news.” Earlier this year, Nadella agreed to a declaration that conversations with chatbots “replace the provision of information on a website on an AI platform without the need to access the underlying information source.”

This kind of language speaks to how technology can directly compete with the original work rather than transform it.

Microsoft’s document says there is a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data based on the training of the underlying model.”

The scale of the copy is astonishing. This document reveals for the first time that OpenAI’s intermediate training dataset alone contains over 91,692 copies of copyrighted works published by the NYT, the Daily News, and the Center for Investigative Reporting. The dataset derived from Common Crawl contained over 2 million documents from nytimes.com alone.

In a January 2023 internal memo, Hecht called it a “stunning theft of unprecedented scale” and “the largest theft of labor in human history.”

The filing provides new details about how OpenAI and Microsoft obtained the plaintiff’s content, including scraping it from Bing Index.

“OpenAI provided the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI’s models within its commercial products,” the filing states. “Microsoft has similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.”

The companies allegedly assembled Project Mango data into a training dataset that included at least 160,903 copies of original works from news publishers.

To get the most out of scraping, OpenAI employees allegedly devised a plan to bypass paywalls without being detected. According to the filing, OpenAI researcher Nick Rider told Brockman about “hacks to get around paywalls at any time,” to which Brockman responded, “Oh, that’s nice.”

OpenAI employees also allegedly built training datasets such as WebText and WebText2 that relied heavily on scraped news content. It is also said to have pulled millions of articles from Common Crawl, a free and open repository of web crawl data. “The findings also explain a deliberate effort to remove copyright notices from the training data before reaching the model, because we don’t want the model to output a ‘copyright notice’ to the user,” the researchers wrote.

OpenAI and Microsoft did not respond to requests for comment.

If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.



Source link

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Editor-In-Chief
  • Website

Related Posts

Even the British King is hesitant about AI.

September 17, 2026

UN turns to Google to make global data available to AI agents

September 17, 2026

2 days left until exhibiting at TechCrunch Disrupt 2026

September 17, 2026
Add A Comment

Comments are closed.

News

Can the “middle powers” ​​come together to build a new economic union? |European Union News

By Editor-In-ChiefSeptember 17, 2026

Canada’s Prime Minister outlines plans to strengthen relations with the EU.Canada’s prime minister is turning…

Canada’s Mr. Carney welcomes EU membership proposal | European Union News

September 17, 2026

India warns new US tariffs on Russian oil could impact relations | Indian Oil and Gas News

September 17, 2026
Top Trending

Even the British King is hesitant about AI.

By Editor-In-ChiefSeptember 17, 2026

King Charles hosted a private summit on Thursday with some of the…

New unedited filing reveals Microsoft executive calls AI scraping ‘the greatest labor theft in human history’

By Editor-In-ChiefSeptember 17, 2026

New unredacted information in a copyright lawsuit filed by The New York…

UN turns to Google to make global data available to AI agents

By Editor-In-ChiefSeptember 17, 2026

The United Nations announced Thursday that it is working with Google to…

Subscribe to News

Subscribe to our newsletter and never miss our latest news

Welcome to WhistleBuzz.com (“we,” “our,” or “us”). Your privacy is important to us. This Privacy Policy explains how we collect, use, disclose, and safeguard your information when you visit our website https://whistlebuzz.com/ (the “Site”). Please read this policy carefully to understand our views and practices regarding your personal data and how we will treat it.

Facebook X (Twitter) Instagram Pinterest YouTube

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Facebook X (Twitter) Instagram Pinterest
  • Home
  • Advertise With Us
  • Contact US
  • DMCA Policy
  • Privacy Policy
  • Terms & Conditions
  • About US
© 2026 whistlebuzz. Designed by whistlebuzz.

Type above and press Enter to search. Press Esc to cancel.