Robot humanoid reading a blue book against a light blue background.

Can AI Be Trained on Copyrighted Books? Courts Are Split, and the Stakes for ChatGPT and Rivals Are Huge

AI copyright law is unsettled as courts weigh whether training models on books is fair use, piracy or market competition.

Updated August 23, 2026 6:26 pm

In short

U.S. courts remain split on AI training and copyright, with legality often turning on whether the data was pirated and whether the resulting product directly competes with the original work.

  • Courts are treating AI training and piracy as separate legal questions.
  • Anthropic’s case suggests illegal data sourcing can trigger major liability even if training itself may be lawful.
  • The Thomson Reuters v. Ross Intelligence ruling shows direct market competition weakens fair use defenses.
  • A fully AI-generated work is not copyrightable under current U.S. law, adding another layer of uncertainty.
  • Congress has not updated copyright law since 1976, leaving judges to improvise case by case.

Update — August 23, 2026 6:26 pm

The updated source adds that copyright law has not been revised since 1976, which helps explain why judges are using decades-old rules to sort out AI-training disputes.

It also sharpens the contrast with the Ross Intelligence case, where a judge said training on Reuters material to build a rival legal product was not fair use because it was not meaningfully different from the original.

The new version further notes that courts have not yet accepted the argument that chatbot outputs compete with authors by generating synthetic books, and that the issue will likely stay unsettled until later stages of ongoing lawsuits.

U.S. courts are still deciding whether AI companies can legally train models on copyrighted books, and the answer is shaping up to be highly dependent on how the material was obtained and how the models are used. The issue matters because rulings could determine whether companies behind ChatGPT, Gemini, Claude and other systems can keep using vast libraries of books and articles without paying authors.

At the center of the debate is a basic but unresolved question: does reading, ingesting and statistically learning from a copyrighted work count as infringement, or is it a lawful transformation of the original material? Recent decisions suggest judges are willing to separate “training” from outright piracy, but they are also drawing a harder line when companies use copied works from illegal repositories or when the new product is seen as a direct market substitute.

The legal uncertainty has turned copyright law into one of the most important pressure points in the generative AI boom. Authors, publishers and media companies say their work has been used without consent; AI firms argue that model training is closer to learning than copying. For now, the law is being made case by case, and neither side can claim a definitive win.

Why the question is suddenly so important

The fast rise of generative AI has forced courts to apply older copyright rules to a technology Congress never specifically addressed. That gap is now at the center of disputes involving book publishers, news organizations, legal databases and AI companies that train on enormous datasets scraped from the internet.

Much of the tension comes from scale. Modern models are built on massive collections of books, papers, websites and other texts, often gathered in ways that authors never approved. That has left writers and publishers arguing that their livelihoods are being undercut by systems trained on their work. AI firms counter that training a model is not the same thing as republishing the underlying books.

Copyright and AI are colliding in an area of law that is still unsettled, with lawyers and judges trying to fit a new technology into rules written decades ago, according to intellectual-property attorney Cathy Gellis.

Gellis said the issue is unusually difficult because both the technology and the emotional stakes are intense, with strong arguments on each side. That complexity is part of why the legal landscape remains fragmented.

What did the Anthropic ruling actually say?

The Anthropic case became one of the most closely watched AI copyright decisions when a federal judge ordered the company to pay $1.5 billion to settle claims brought by authors whose books were used in training. But the ruling was more nuanced than a simple victory for writers.

Judge William Alsup did not hold that AI training itself was unlawful. Instead, he found fault with the company’s use of books that had been taken from illegal shadow libraries. In other words, the problem was the source of the books, not necessarily the act of training a model on them.

In the judge’s view, the model was functioning more like a reader studying texts and creating something new than a machine reproducing the original books verbatim.

That distinction matters because it gives AI developers a legal argument that training can be transformative. It also gives authors a separate argument that copying books from unauthorized libraries is still infringement, even if the downstream training process is treated differently.

How the case may help AI companies

The Anthropic ruling is being read by some lawyers as broadly favorable to model builders because it suggests that courts may distinguish between the acquisition of data and the learning process itself. Cathy Gellis said the decision could be seen as relatively good news for AI firms because it treated model training more like reading than copying.

That matters in copyright law, which is centered on unauthorized reproduction. If a court sees model training as analogous to a person studying books and internalizing patterns, companies may be able to rely on fair use arguments. If a court sees training as making impermissible copies for commercial exploitation, the risks become much higher.

Even so, the size of the Anthropic penalty highlights how expensive these disputes can become. A $1.5 billion settlement is enormous in any context, but especially if a company expects to grow into a business worth hundreds of billions of dollars.

Case / issue What happened Why it matters
Anthropic copyright case Judge approved a $1.5 billion settlement over books taken from illegal shadow libraries. Suggests the source of training data can be decisive, even if model training itself is not banned.
Thomson Reuters v. Ross Intelligence A court found training on Reuters material to build a competing legal tool was not fair use. Shows that direct market competition can weaken an AI company’s legal defense.
Thaler v. Perlmutter A fully AI-generated work was ruled not copyrightable. Raises new questions about authorship, proof and how courts should treat human-AI collaboration.

How does fair use apply to AI training?

Fair use is the doctrine most AI companies are relying on, but it is not a blanket permission slip. It allows limited use of copyrighted material without authorization in certain situations such as criticism, teaching, parody and commentary. Courts weigh several factors, including the purpose of the use, the nature of the work, how much was taken and whether the new use harms the original market.

In AI disputes, the most important factor is often whether the use is “transformative.” That means judges ask whether the new product serves a meaningfully different purpose from the original. If the answer is yes, the defendant has a stronger fair use case. If the answer is no, the claim weakens considerably.

Jason Henderson, a senior attorney and founder of JWL International’s IP & Media practice, said the courts are signaling that competitive overlap may be the deciding issue in many cases. In his view, companies using other people’s property to build tools that directly compete with the original creators face a much harder road.

Henderson said courts appear more skeptical when the new AI product is meant to compete with the original work, while uses that do not directly substitute for the source material are more likely to survive legal scrutiny.

That logic helps explain why some disputes are going in different directions. In one case, judges have been more receptive to arguments that training is transformative. In another, they have focused on whether the resulting product would compete in the same market as the copyrighted material.

Why market competition is the key battleground

Copyright law is not only about ownership; it is also about preserving markets for creators. That is why courts often ask whether the allegedly infringing use substitutes for the original or siphons away revenue.

For authors, the concern is that a chatbot can absorb the style, structure and substance of thousands of books and then generate outputs that reduce demand for the originals. For AI companies, the counterargument is that a model does not resell a book, and its outputs are generated through pattern recognition rather than replication.

The dispute becomes even sharper when AI systems are built to perform the same function as the original work. If a model trained on legal content is marketed as a replacement for a legal research platform, for example, courts may see that as a competitive injury rather than a transformative use.

What happened in the Thomson Reuters case?

The Thomson Reuters case against Ross Intelligence is one of the clearest examples of a court treating AI training as unfair when the goal is direct competition. Ross allegedly copied Reuters material to build an AI-powered legal research product that would challenge Thomson Reuters’ own platform.

Judge Stephanos Bibas concluded that the use was not transformative because it did not serve a meaningfully different purpose from the original legal research material. That decision gave publishers and rights holders a stronger foothold in the argument that copying content to make a substitute product is not protected simply because AI is involved.

This is the line many copyright owners want courts to draw more broadly. They argue that if companies can ingest protected works to create rivals, the market for books, journalism, legal research and other licensed content could be hollowed out.

AI companies, meanwhile, point out that not every model built on copyrighted material is intended to replace the source. A general-purpose chatbot, they argue, is not the same thing as a legal database or a single book, even if both rely on text from elsewhere.

How is training different from copyrighting AI-generated work?

Training a model and claiming copyright in its output are related but separate legal questions. The first asks whether a company can use copyrighted works to build an AI system. The second asks whether a work created entirely by AI can itself be protected by copyright.

In Thaler v. Perlmutter, a court ruled that a work generated entirely by AI does not qualify for copyright protection under current law. That decision underscores a foundational principle of U.S. copyright: protection generally requires human authorship.

This creates a messy gray area for creators and companies using AI tools in mixed workflows. If a human writes a book with AI assistance, how much human contribution is enough? If a work is edited, polished or substantially altered by a person, does that make it copyrightable? Courts have not given a bright-line answer.

Gellis compared the issue to using ordinary software tools such as Microsoft Word, noting that most people would not think a spelling checker gives the software company ownership over the manuscript.

That analogy helps explain why many lawyers think AI is forcing courts to revisit assumptions that were easy to ignore before generative systems became capable of producing polished text, images and code on their own.

Why the law feels so unsettled

The central reason is simple: the U.S. Copyright Act has not been comprehensively updated since 1976. Judges are being asked to interpret rules drafted long before the internet, web scraping, cloud computing or large language models existed.

As a result, judges are developing the law in fragments. One court may view model training as a productive form of learning. Another may focus on whether the dataset was lawfully obtained. A third may emphasize whether the output competes with the copyrighted work. The result is a patchwork of rulings that can be difficult for companies and creators to navigate.

Henderson said the uncertainty is contributing to widespread anxiety because everyone involved understands how much AI systems depend on huge quantities of text. Until appellate courts or Congress provide more clarity, the basic question remains unresolved.

What judges are likely to focus on next

Several recurring themes are already emerging across cases:

  • How the training data was obtained: lawful licensing may matter more than ever.
  • Whether the new product substitutes for the original: direct competition is legally dangerous.
  • How much of the original work is used: scale and access can affect the fair use analysis.
  • Whether the output reproduces protected expression: outputs that mirror source material are more vulnerable.
  • Who can claim authorship: mixed human-AI works create new proof problems.

Could authors argue that chatbots compete with books?

Yes, and that argument is likely to be tested more aggressively in future litigation, but it has not yet carried the day in a broad, final way. Authors may claim that chatbots trained on novels or nonfiction can produce summaries, stylistic imitations or even synthetic books that reduce demand for the originals.

That theory is important because it moves the fight beyond the question of copying and toward market substitution. If a chatbot can answer questions, mimic a writer’s style or generate book-like text on demand, plaintiffs may argue it is competing with the market for books in a way that fair use should not protect.

AI companies will respond that such systems are still tools for generating new text, not replacements for specific books. They will also argue that broad training is necessary to create general-purpose models and that blocking access to training data could slow innovation across the entire field.

For now, no court has issued the final word on whether model training on books is categorically legal or illegal. Instead, the law is being developed through a series of narrower findings about piracy, market competition, transformation and authorship.

What happens next?

The next phase of the fight will likely come through more litigation, possible appeals and potentially new legislation. Because these cases are still working their way through the courts, the industry does not have a final answer yet.

That means AI companies are operating in an environment where early rulings can shape licensing strategy, data sourcing and product design. Some firms may be more careful about where they get their training material. Others may continue to defend broad fair use positions while the courts continue sorting out the issue.

For writers and publishers, the stakes are equally high. A ruling that broadly protects training could make it harder to force AI companies into licensing deals. A ruling that narrows fair use could strengthen authors’ leverage and push the industry toward paid access to copyrighted works.

What is clear is that the legal system is still trying to answer a question that goes to the heart of generative AI: when does learning from human creativity become copying it? Until courts settle that boundary, the battle over copyrighted books will remain one of the defining legal fights of the AI era.

Key issue Authors’ view AI companies’ view
Training on books Unauthorized use of protected works Equivalent to learning from reading
Illegal source material Always problematic and compensable A separate issue from training itself
Fair use Limited when outputs compete with originals Protects transformative model training
AI-generated works Need clearer rules for mixed authorship Current law should adapt to new tools

Frequently asked questions

Is it legal to train AI models on copyrighted books?

It depends. U.S. courts have not issued one universal rule, but recent decisions suggest that training may be treated as fair use in some situations, especially when the use is transformative, while copying books from illegal sources can still create liability.

Why did Anthropic owe $1.5 billion if the training was ruled lawful?

Because the court focused on how Anthropic obtained the books, not just on training itself. The liability came from using pirated copies taken from shadow libraries, which the judge treated as unlawful even though the model-training theory was not automatically banned.

What does fair use mean for AI companies?

Fair use can protect some uses of copyrighted material without permission, but only if the use is sufficiently transformative and does not unfairly damage the original market. For AI firms, that usually means courts will ask whether the model is competing with the copyrighted work or doing something meaningfully different.

Can a fully AI-generated work be copyrighted?

No, not under current U.S. law. A court ruled in Thaler v. Perlmutter that a work created entirely by AI lacks the human authorship needed for copyright protection, although mixed human-AI works remain a gray area.

Why are publishers and authors worried about AI training?

They fear that AI systems trained on their books and articles can imitate style, generate substitutes and reduce demand for original works. That could weaken sales and licensing revenue unless courts or lawmakers require permission or compensation.

Share this 🚀