In short
New unsealed filings in The New York Times’ lawsuit against OpenAI and Microsoft allege mass AI scraping of news content, paywall bypass and internal warnings about publisher harm. The documents could strengthen arguments that the companies knew their products might compete directly with the original journalism.
- The Times’ latest filing alleges OpenAI and Microsoft used scraped news content at massive scale.
- Internal statements quoted in the brief describe AI training as theft and warn of existential harm to publishers.
- Microsoft’s Copilot and OpenAI chatbots are alleged to have reduced traffic to original news sites.
- The case is becoming a major test of whether AI training on copyrighted journalism qualifies as fair use.
Newly unsealed material in The New York Times’ copyright case against OpenAI and Microsoft says senior executives privately described AI training on news content as theft, while internal documents acknowledged that chatbot products could undermine publishers’ traffic and business models. The filings matter because they add fresh evidence to one of the most important legal fights in generative AI: whether companies can train models on copyrighted journalism without permission or payment.
The latest disclosure does not end that dispute, but it sharpens the stakes. According to the court filing, Microsoft and OpenAI allegedly gathered news content at massive scale, removed copyright notices from training data, and used methods designed to avoid paywalls. The documents also suggest both companies understood that their tools could substitute for the original sources that created the work in the first place.
Much of the newly public information comes through The Times’ own legal brief rather than the underlying exhibits, which remain under seal. Still, the material offers a rare look at how some of the biggest names in AI internally discussed the commercial and legal risks of building products on scraped journalism.
What the new filings say about AI training and news content
The unredacted filing expands a lawsuit that The New York Times brought against OpenAI and Microsoft three years ago, when it accused the companies of training generative AI systems on its reporting without authorization. The new material alleges that the companies did not merely ingest publicly available material in a passive way. Instead, the filings say they built large-scale datasets from scraped web content, pulled articles from search indexes and free crawl repositories, and in some cases attempted to bypass paywalls.
The central legal issue remains whether that training qualifies as fair use under U.S. copyright law. Courts have generally been receptive to AI companies’ argument that model training can transform protected works rather than replace them. But the newly disclosed statements could complicate that defense, especially where internal remarks suggest the outputs were expected to compete directly with the original publishers.
The Times’ latest brief, according to the filing, argues that these admissions show the firms understood the difference between transformative use and market substitution. That matters because fair use typically depends not just on how a work is copied, but also on whether the use harms the market for the original.
Why do the filings matter now?
They matter because they could influence how courts think about AI training, licensing and publisher compensation at a moment when the legal landscape is still unsettled. The case arrives as lawmakers, courts and the White House debate how generative AI should interact with copyright. Earlier this month, the Trump administration filed a brief supporting OpenAI’s position that unlicensed use of copyrighted material for training can be lawful in some circumstances.
But the new evidence gives the Times fresh ammunition. The documents, as described in the filing, suggest that Microsoft’s and OpenAI’s own people were aware of the commercial damage that could come from training on news without paying for it. If a court finds that the companies knew their products could divert traffic away from the original publishers, that could weaken a core argument that the training was legally protected because it was transformative.
Below is a quick summary of the main allegations disclosed in the latest filing.
| Issue | What the filing alleges | Why it matters |
|---|---|---|
| Training data collection | Mass scraping, use of Bing Index material and Common Crawl, plus data sharing between Microsoft and OpenAI | Shows the scale and methods of acquisition |
| Paywall circumvention | Employees allegedly discussed a “hack” to get around The Times’ paywall | Could undermine claims of innocent or incidental use |
| Copyright notices | Training data allegedly had copyright notices removed before model ingestion | Suggests deliberate preparation of data for model training |
| Market harm | Internal documents reportedly predicted severe drops in click-through rates | Relevant to the fair-use analysis |
| Scale of copying | Millions of documents and more than 90,000 copies of works in mid-training datasets | Illustrates the breadth of alleged copying |
How did Microsoft and OpenAI allegedly use the content?
They allegedly used it to build training datasets and commercial products, including Microsoft’s Copilot answer engine and OpenAI’s language models. The filing says OpenAI delivered its GPT-3 training set to Microsoft, which then used it to study how to integrate OpenAI models into its own products. Microsoft also allegedly passed data to OpenAI through internal efforts known as Project Taxi and Project Mango.
According to the filing, Project Mango in particular became a source of training material that included copies of at least 160,903 unique works from news organizations. The brief also says a Common Crawl-based dataset contained more than 2 million documents from nytimes.com alone. Separate mid-training datasets allegedly included more than 91,692 copies of works published by The Times, the Daily News and the Center for Investigative Reporting.
The numbers are significant because they suggest the issue was not limited to isolated examples. If the allegations are proven, they would point to an industrial-scale pipeline in which news content was gathered, processed and reused across multiple model-development stages.
What are Project Taxi and Project Mango?
They appear to have been internal Microsoft-OpenAI data-sharing efforts used to support model development and evaluation. The filing says Microsoft and OpenAI exchanged training data through those initiatives, though the exact technical details remain limited because the underlying exhibits are still sealed.
That secrecy is important. Without the underlying records, the public is relying on summaries and quotations contained in the legal brief. Even so, the allegations describe a structured and recurring flow of content between the companies, not an incidental overlap in data sources.
What did Microsoft executives allegedly say?
They allegedly used unusually blunt language to describe the economics of AI training on news content. In one January 2023 internal memo attributed to Microsoft’s Director of Applied Science, Brent Hecht, the company reportedly called the practice “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
Hecht also reportedly warned of a “doom loop” in which lower traffic to publishers would eventually hurt the quality and diversity of the web itself, which in turn would reduce the performance of the models that depend on it.
According to the filing, Microsoft’s internal analysis described it as unusual for a product to threaten the economic base of the suppliers that make the product possible, but said that was the situation the company had created for its LLM business.
That framing is striking because it portrays the relationship between AI firms and publishers as symbiotic but fragile. AI models need fresh, high-quality content to improve; publishers need traffic and subscription revenue to fund that content. If one side undermines the other, the entire ecosystem can deteriorate.
How did the companies view paywalled journalism?
They allegedly treated paywalled journalism as content that should be licensed, not scraped. Microsoft CEO Satya Nadella testified in deposition that anything behind a paywall should be licensed by anyone who wants to use it for grounding or training. He also said that if he had known OpenAI trained on paywalled material, he would have required the company to retrain.
That testimony is particularly relevant because it appears to acknowledge a distinction between publicly visible web content and content that publishers intentionally restrict. In practice, paywalls are one of the main ways news organizations monetize reporting. If AI systems can ingest that content without permission while simultaneously providing answers that reduce site visits, the publisher’s revenue model is under pressure from both directions.
The filings also quote internal OpenAI communications in which the company’s head of ChatGPT, Nick Turley, said publishers faced an existential threat from chatbot products that are “largely substitutive” and likely to become even more substitutive over time. OpenAI President Greg Brockman is also described as saying the models were “excellent at news.” Nadella, according to the filing, testified that chatting with AI can replace the need to click through to an original website for information.
Why the wording is important
The wording matters because fair use often turns on substitution. If a product simply helps users discover or interpret information, courts may view it differently than if the product replaces the original source altogether.
That distinction is central to the Times’ argument. The more the companies’ own documents describe the tools as substitutive, the more the case begins to look less like a transformative research project and more like a competing news product built on copied content.
How much content was allegedly copied?
Enough, according to the filing, to show the alleged copying was systematic rather than accidental. The brief says one dataset included more than 2 million documents from nytimes.com and that OpenAI’s mid-training datasets contained over 91,692 copies of works from The Times, the Daily News and the Center for Investigative Reporting. Another dataset built from Project Mango allegedly included copies of more than 160,000 unique works from news publishers.
Large language models are typically trained on enormous collections of text, so the size of the datasets alone is not unusual in the industry. What matters here is the source mix. A dataset that repeatedly contains news articles, including paywalled material, can raise both copyright and competition concerns, especially if the same articles are later reflected in model outputs or used to improve commercial products that answer users directly.
The filing also says OpenAI and Microsoft used Common Crawl, a widely used public web archive, as a major source. That is a common industry practice. But the Times alleges the companies relied on it in ways that disproportionately captured news reporting and then combined it with other datasets that were more directly tied to the publishers’ work.
What about the alleged paywall bypass?
The filing says OpenAI employees discussed a workaround to get around The Times’ paywall and that Brockman responded positively when told about it. The brief presents that exchange as evidence that the company was not merely scraping whatever was openly available, but actively trying to access restricted content without detection.
If proven, that would be legally and reputationally significant. A company can argue that it relied on publicly accessible information and believed the training process was lawful. It is much harder to defend a practice that allegedly involved known efforts to circumvent access controls.
Still, the legal weight of the allegation will depend on the underlying exhibits, which have not been made public. For now, the filing supplies only a summary of those communications. Courts will eventually have to decide how much to rely on those descriptions and how much evidentiary value to assign them.
How strong is the fair-use defense?
It is still a real defense, but the new material gives critics more to work with. Fair use in the United States considers four factors: the purpose of the use, the nature of the copyrighted work, the amount taken, and the effect on the market for the original.
AI companies tend to argue that training is transformative because the model does not reproduce specific articles in a traditional sense. Publishers respond that even if the training process transforms the material, the resulting system can still compete with the original source by generating answers that displace traffic and subscriptions.
In this case, the newly disclosed statements about substitutive behavior and traffic decline go directly to the fourth factor, market effect. The filing also points to the amount and character of the copying, which could matter under the other factors as well.
The four fair-use factors at a glance
- Purpose and character: Is the use transformative or commercially substitutive?
- Nature of the work: News reporting is generally creative and valuable, which can strengthen protection.
- Amount used: Massive copying weighs against a broad fair-use claim.
- Market impact: If AI tools reduce traffic, subscriptions or licensing value, that is a major concern.
What does this mean for publishers?
It means the fight over licensing is likely to intensify. News organizations have already been seeking deals that pay for the use of their work in AI training and product outputs. The filings support the argument that publishers are not only asking for compensation out of principle, but because they believe their content is becoming infrastructure for products that can cannibalize the very audience those publishers depend on.
Some publishers see AI partnerships as necessary and potentially lucrative. Others fear they will become invisible labor sources whose reporting is harvested to train systems that then summarize, repackage or replace their work. The new filing lends force to the second concern.
There is also an industry-wide concern here beyond The New York Times. If one of the world’s best-funded AI companies is alleged to have used paywalled journalism without permission, the implications extend to thousands of smaller outlets that have even less leverage to negotiate contracts.
What happens next in the lawsuit?
The case will continue to move through discovery, motion practice and likely additional disclosure battles. The unsealed brief may prompt more arguments over which internal records should remain confidential and which should be made public. It may also shape how judges view requests related to model data, datasets and product design documents.
For OpenAI and Microsoft, the immediate challenge is reputational as well as legal. The public release of internal comments using words like theft and existential threat makes the companies’ messaging more difficult. Even if they ultimately win on fair use, they now face a record that can be cited by critics, lawmakers and rivals whenever questions arise about how AI systems were built.
For publishers, the filings could become a blueprint for future litigation and licensing negotiations. They offer a vocabulary for describing harm, a map of alleged data flows and a set of internal statements that may resonate with judges who are trying to decide whether generative AI is a new kind of tool or simply a more powerful form of market substitution.
Timeline of the dispute
| Date | Event | Why it matters |
|---|---|---|
| 2023 | The New York Times files its copyright lawsuit against OpenAI and Microsoft | Launches one of the most closely watched AI copyright cases |
| January 2023 | Microsoft internal memo reportedly calls the practice unprecedented theft | Shows early internal concern about the ethics and economics of training |
| January 2024 | Microsoft presentation warns of a “doom loop” tied to search and traffic loss | Links AI product design to publisher harm |
| Earlier in 2026 | Satya Nadella gives deposition testimony on paywalled content | Suggests license requirements for restricted material |
| September 17, 2026 | New unredacted filing reveals additional internal statements and alleged data practices | Escalates the dispute with new detail on training and market impact |
Who is at the center of the case?
At the center are two of the most influential companies in generative AI and one of the world’s most prominent newspapers. OpenAI is the developer behind ChatGPT and related large language models. Microsoft is both a major OpenAI investor and a commercial partner that has embedded OpenAI models into products such as Copilot. The New York Times is the plaintiff, arguing that its journalism was used to power those systems without permission.
That combination makes the lawsuit important far beyond the parties involved. A ruling that favors The Times could push the industry toward broader licensing agreements. A ruling that favors OpenAI and Microsoft could strengthen the legal case for training on large collections of copyrighted material, at least under current U.S. law.
Either way, the filings ensure that the debate over AI and copyright is no longer abstract. It is now tied to specific words, specific datasets and specific business consequences.
Internal remarks cited in the filing suggest company leaders understood that the products could both rely on publishers’ work and threaten the publishers’ business model at the same time.
That tension may ultimately be what the case is about: whether the AI industry can build the next generation of information products without undercutting the institutions that produce the information itself.
For now, the new material does not decide the legal issue. But it does make one thing clearer: the dispute is no longer just about whether AI models were trained on news. It is also about what the companies knew, what they said internally, and whether they built their products with full awareness that the result might be, in the words attributed to their own executives, economically devastating for the sources they depended on.
Frequently asked questions
What do the new filings in The New York Times case allege?
The new filings allege that OpenAI and Microsoft scraped news content at scale, bypassed paywalls, stripped copyright notices from training data and used that material to build AI models and products that could compete with publishers.
Why are the internal quotes in the filing important?
The internal quotes are important because they suggest company leaders knew the products could harm publishers economically. That could weaken a fair-use defense, which depends in part on whether the use substitutes for the original work and damages its market.
Did Microsoft or OpenAI respond to the allegations?
The companies did not immediately respond to requests for comment in the source material. The filing itself is part of ongoing litigation, so the companies will likely continue to contest the claims in court.
How does this lawsuit affect the AI industry?
It could influence how courts and regulators treat AI training on copyrighted material. A publisher victory could accelerate licensing deals, while an AI-company victory could make it easier to train models on large copyrighted datasets without direct permission.









