In short
Microsoft is arguing that Copilot rarely reproduces substantial text from books and news stories, using millions of chat logs to support its fair-use defense in a copyright case brought by publishers and authors. The outcome could influence how courts treat AI training on copyrighted material.
- Microsoft says fewer than 1% of 8.2 million Copilot chats contained 16 or more matching words from news content.
- The company is using the numbers to argue that AI training on copyrighted works is transformative and qualifies as fair use.
- Publishers and authors claim Copilot and related AI products can reproduce protected text and compete with original works.
- The dispute is part of a broader copyright battle involving Microsoft, OpenAI, The New York Times and book authors.
- A judge’s ruling on summary judgment could determine whether the case ends early or moves toward trial.
Microsoft says its Copilot chatbot almost never reproduces substantial passages from news articles or books, arguing that the rare instances of overlap are too limited to support publishers’ and authors’ claims that the company’s AI tools are built on copyright infringement. The filing, made in federal court on Friday, is part of Microsoft’s push to end the case early and keep the broader debate over AI training on its side of the fair-use line.
The company’s latest legal argument matters because the dispute is not just about Microsoft and OpenAI; it is about whether the way large language models are trained and deployed can legally rely on copyrighted text at scale. If judges accept Microsoft’s view, the decision could shape the future of AI development, newsroom licensing strategies, and the economics of book and media publishing.
In its court papers, Microsoft says an analysis of millions of Copilot conversations showed only tiny levels of direct text overlap with protected works. The plaintiffs, which include The New York Times, other publishers, and book authors, say the products are built using their material and can compete with the original works by reproducing copyrighted text on demand.
What Microsoft is telling the court
Microsoft is using new discovery material to argue that Copilot does not typically spit out long, recognizable passages from copyrighted sources. According to the company, a review of 8.2 million chat logs found that fewer than 1% of conversations included at least 16 words matching news content used to ground the model.
Microsoft says the logs were not a random slice of user activity. The company says they were selected because they contained keywords tied to publishers’ websites, making them more likely to surface material from the news organizations involved in the suit. Even with that supposedly favorable sampling, Microsoft says only a small fraction of chats showed meaningful overlap.
Microsoft’s legal position is that the low rate of verbatim reproduction supports its view that AI training is transformative and should be treated as fair use, even when copyrighted works are included in the training pipeline.
The company also says an expert for the Center for Investigative Reporting identified just 51 examples of “substantial overlap” in the dataset, while an expert in the authors’ case found only 24 Copilot responses with 30 or more matching words across the 8.2 million conversations. Microsoft further claims that, of 212 books examined, only 10 showed any matches.
How Microsoft is framing fair use
Microsoft is not denying that copyrighted material plays a role in training modern AI systems. Instead, it is asking the court to view that use through the lens of transformation: the idea that a model trained on books, articles, and other text creates something fundamentally different from the underlying works.
That argument has become central to the industry’s defense in multiple lawsuits. AI companies contend that training a model is more like learning patterns, structure, and language than copying a book or article for readers to consume. Publishers and writers counter that the models are commercially valuable precisely because they ingest large amounts of protected text and can sometimes reproduce it.
Microsoft says occasional regurgitation does not defeat the broader purpose of model training. In the company’s view, rare verbatim outputs are not evidence that the system is designed to replace the original works. Rather, the company argues, such examples are edge cases that do not alter the legal character of the technology.
Why the overlap numbers matter
The numbers are important because they may help a judge decide whether the plaintiffs can show harm substantial enough to survive a motion for summary judgment. If Microsoft convinces the court that Copilot usually does not reproduce protected text in a way that substitutes for the original, the company could narrow the case before trial.
For publishers and authors, the same data cuts the other way. They argue that even a small number of direct reproductions can be evidence of systematic misuse, especially if the model can be prompted to echo valuable or timely content. In copyright law, the dispute is not only about frequency but also about context, purpose, and market impact.
What the lawsuit is really about
The Microsoft filing is part of a larger legal fight brought by news organizations and book authors against Microsoft and OpenAI. The plaintiffs say the companies trained AI systems on their works without permission and then released products that compete with the original sources by summarizing, paraphrasing, and sometimes echoing copyrighted text.
The cases were consolidated under one judge, a move intended to streamline discovery and motion practice. The publishers and authors objected to that consolidation, but the court pressed ahead anyway. Microsoft is now asking for summary judgment, a ruling that would end the case without a full trial if the judge concludes the legal record is already clear enough.
OpenAI is also a central defendant in the broader dispute, but Microsoft’s role is especially significant because Copilot is one of the company’s most visible consumer and enterprise AI products. The outcome could influence not only chatbot behavior but also how closely AI vendors need to police output for potential regurgitation.
| Issue | Microsoft’s position | Publishers’ and authors’ position |
|---|---|---|
| AI training on copyrighted works | Permissible fair use because the model is transformative | Unlicensed copying to train commercial tools is infringement |
| Copilot output | Rarely reproduces significant text from protected sources | Can still echo copyrighted passages and substitute for originals |
| Scale of overlap in logs | Fewer than 1% of 8.2 million chats had 16+ matching words | Even limited repeats may show systemic misuse |
| Legal goal | Early dismissal via summary judgment | Continue toward trial and damages |
How did Microsoft assemble its evidence?
Microsoft says it turned over 8.2 million Copilot chat logs during discovery for review by an expert retained by the news publishers. The company claims those logs were deliberately selected because they were most likely to surface use of the plaintiffs’ material, based on keyword signals associated with their sites.
That detail is critical. In legal fights over AI training and output, plaintiffs often argue that a defendant’s product can generate harmful examples even if those examples appear infrequently in a broad sample. Defendants, by contrast, try to show that the rare instances are not representative of ordinary use. Microsoft is clearly trying to position this dataset as a worst-case sample that still yielded very little direct text reuse.
The company also points to separate expert reviews in the authors’ and publishers’ cases. It says those analyses found very few matches in both the news and book materials evaluated. Although Microsoft did not release the full underlying data in its filing, it is using the cited expert findings to argue that the plaintiffs cannot prove widespread copying through Copilot outputs.
Who is missing from the record?
The New York Times, the Center for Investigative Reporting, and the Authors Guild did not immediately respond to requests for comment on Microsoft’s filing. That leaves the company’s account unchallenged for now in the public record, though the plaintiffs are expected to answer in court filings of their own.
Because the case is still in the motion stage, both sides are likely to continue framing the evidence in the most favorable way possible. Microsoft wants the judge to see the numbers as proof of limited overlap; the plaintiffs will likely argue that the same numbers reveal that the company was able to locate, test, and quantify the unauthorized use of their works in a commercial AI system.
Why this case could set an industry precedent
This lawsuit has implications far beyond a single chatbot. If Microsoft and OpenAI persuade the court that training on copyrighted material is fair use and that occasional reproductions are legally tolerable, other AI companies could point to the decision as cover for their own training practices.
If the plaintiffs succeed, the result could force a shift toward licensing deals, tighter output controls, and perhaps a more expensive path for building general-purpose AI systems. That could benefit publishers and authors who want payment for the use of their work, while increasing costs for AI vendors and potentially reducing the amount of content they can train on freely.
The case also arrives in a political environment where federal agencies and the White House have been paying close attention to AI copyright questions. A filing from the Trump administration in the separate New York Times matter this week supported OpenAI, underscoring that the legal and policy battles around AI training are now running in parallel with the courtroom fight.
What happens next?
The immediate question is whether the judge will grant Microsoft’s request for summary judgment. If the court agrees, the case could end before a jury ever hears it. If the judge rejects the motion, the plaintiffs will continue toward trial and the parties will keep fighting over how copyright law applies to AI systems built on large-scale text ingestion.
Either way, the dispute is likely to remain a reference point for publishers, authors, and AI developers. The central tension is straightforward: AI companies say they need copyrighted material to build useful systems, while rightsholders say those systems must not be allowed to absorb and echo their work without payment or permission.
For now, Microsoft is betting that the numbers are on its side. The company is telling the court that Copilot’s rare reproductions are too isolated to overcome the argument that its AI model training is transformative, and therefore lawful. The plaintiffs will have to persuade the judge that even limited instances of copying reveal a deeper problem in how the technology works and how it competes with the original sources.
Key figures at a glance
| Metric | Figure | Meaning |
|---|---|---|
| Copilot chat logs reviewed | 8.2 million | Large discovery set used in the dispute |
| Chats with 16+ matching words | 59,545 | Microsoft says this is under 1% of the total |
| Book responses with 30+ matching words | 24 | Microsoft says this is a tiny share of the sample |
| Books with any matches | 10 of 212 | Microsoft says most evaluated books showed no overlap |
| Substantial overlap examples cited by CIR expert | 51 | Microsoft says this remains limited in context |
As the case advances, the dispute will likely test more than Microsoft’s Copilot output. It may also test how much evidence courts require before deciding that training and deploying generative AI on copyrighted works crosses the line from innovation into infringement.
Frequently asked questions
What is Microsoft arguing in the Copilot copyright case?
Microsoft is arguing that Copilot rarely reproduces meaningful chunks of copyrighted news or book text, and that the limited overlap is not enough to defeat its fair-use defense. The company says AI training is transformative because the model serves a different purpose than the original works.
How many Copilot chats did Microsoft review?
Microsoft says it reviewed 8.2 million Copilot chat logs during discovery. The company says the logs were selected because they were likely to surface use of the plaintiffs’ websites and works, making them a targeted sample rather than a random one.
Why do publishers and authors say this matters?
Publishers and authors say it matters because AI products can absorb their work without permission and then compete with them by summarizing or reproducing similar text. They argue that even limited copying can show a broader pattern of infringement and market harm.
What would summary judgment mean in this case?
Summary judgment would mean the judge decides the case without a full trial if the legal record is already sufficient. If Microsoft wins that motion, the case could end early; if the judge rejects it, the plaintiffs would continue toward trial.
Does Microsoft deny that Copilot uses copyrighted material?
No. Microsoft does not deny that copyrighted works are part of modern AI training. Instead, it argues that using those materials to train large language models is lawful fair use because the resulting system is different in purpose and function from the source works.









