In short
Sony Music and Warner Chappell have sued Anthropic, accusing the AI company of using tens of thousands of copyrighted songs and lyrics without permission to train Claude. The publishers say the alleged infringement and removal of copyright data could expose Anthropic to billions in damages.
- Sony Music and Warner Chappell say Anthropic used copyrighted songs and lyrics without permission.
- The suit seeks statutory damages that could reach billions if maximum penalties are awarded.
- The complaint also names Anthropic co-founders Dario Amodei and Benjamin Mann.
- Publishers allege the company used pirated books, BitTorrent and scraped lyric sources.
- The case adds to growing legal pressure on AI firms over training data and licensing.
Sony Music and Warner Chappell have sued Anthropic in federal court in California, accusing the AI company of using tens of thousands of copyrighted song lyrics and compositions without permission to train and run its Claude models. The publishers say the case could expose Anthropic to billions of dollars in damages and could become one of the most consequential copyright fights yet over generative AI.
The complaint, filed in the US District Court for the Northern District of California, says Anthropic and its founders built a commercial AI business on what the publishers describe as large-scale infringement of music rights. The filing is the latest in a widening wave of legal challenges aimed at the data practices behind frontier AI systems, and it lands while Anthropic is already under pressure from other rights holders in the music industry.
At the center of the dispute is a familiar but increasingly urgent question for AI developers: how much copyrighted material can be used to build an AI model before that use becomes unlawful? For the music business, the answer in this case is straightforward. Sony and Warner say Anthropic crossed the line repeatedly, at scale, and for profit.
What Sony and Warner are alleging
Sony Music and Warner Chappell accuse Anthropic of copying protected lyrics, removing copyright information, and relying on pirated materials to support the development of its chatbot platform. Their lawsuit portrays the company’s conduct as systematic rather than accidental, arguing that the alleged misconduct involved massive volumes of works and was tied directly to the commercial value of Claude.
According to the complaint, the plaintiffs are seeking statutory damages that could reach $150,000 for each copyrighted work at issue, plus up to $25,000 for each instance in which copyright management information was allegedly stripped away. If a court accepted the publishers’ full damage theory, the total exposure could climb into the billions of dollars.
The suit also names Anthropic co-founders Dario Amodei and Benjamin Mann as defendants, a move that signals the publishers are not limiting their complaint to the corporate entity alone. By targeting the company’s leadership, the plaintiffs appear to be arguing that the alleged infringement was not simply the result of a rogue dataset or a third-party vendor, but part of a broader operational strategy.
Why this lawsuit matters now
This case matters because it comes at a moment when AI firms are under mounting legal and financial pressure to explain how their models were trained. Copyright owners across publishing, music, media, and other creative industries have been challenging the assumption that scraping or downloading works at internet scale is legally defensible simply because the output is probabilistic or transformative.
Anthropic is already dealing with major legal headwinds. The company recently reached a $1.5 billion settlement in a separate lawsuit brought by book publishers, and it has also been sued by Universal Music Group, Concord, ABKCO, BMG, and Round Hill Music. The new complaint suggests that the music industry is not backing away from litigation, even as some AI companies have begun settling certain disputes.
Because the music industry depends heavily on licensing, collective rights management, and metadata tracking, the allegations involving copyright data removal are especially sensitive. If a company is found to have stripped identifying information from song files or lyric records, that can carry separate legal consequences beyond the underlying copying claim.
How the publishers say Anthropic got the material
The publishers’ account says Anthropic’s data collection went far beyond casual web scraping. One of the most striking claims is that co-founder Benjamin Mann allegedly used BitTorrent to download more than five million pirated books, while company employees allegedly downloaded at least two million additional books from a site known as Pirate Library Mirror.
The complaint also says Anthropic scraped lyrics from services such as MusixMatch and LyricFind, both of which license material from rights holders. That distinction matters because it suggests the company allegedly accessed content not merely from public webpages, but from licensed distribution channels whose business model depends on paid permission.
In addition to the bulk-copying allegations, Sony and Warner say some of the songs identified in Anthropic’s training materials include well-known works such as Marvin Gaye and Tammi Terrell’s “Ain’t No Mountain High Enough,” Bon Jovi’s “Livin’ On a Prayer,” Earth, Wind & Fire’s “September,” Leonard Cohen’s “Hallelujah,” and Taylor Swift’s “Paper Rings.”
The publishers argue that Anthropic and its founders carried out a wide-ranging campaign of “illegally torrenting, scraping, and downloading copyrighted works” to build and profit from Claude.
Anthropic did not immediately respond to a request for comment, leaving the allegations unanswered publicly at the time of filing. As in many AI copyright suits, the company is likely to argue that its data use was lawful, that the claims are overbroad, or that the case misunderstands how model training works. But those defenses have yet to be tested in this specific music-rights context.
What songs and rights are at issue?
The complaint identifies both the text of lyrics and the broader rights ecosystem around songs as core parts of the dispute. That is important because music rights are layered: composition rights can be controlled by publishers, while sound recording rights usually belong to labels or performers. This lawsuit focuses on the publishing side, where lyric ownership and synchronization rights are central sources of value.
For AI systems, lyrics can be useful training material because they contain structured language, style, rhyme, repetition, and cultural references. For rights holders, however, lyric databases are exactly the kind of material that is licensed for a fee and protected through contract and copyright law. If a model can reproduce or imitate lyrics, it can threaten the market for those licensed services.
Here is a summary of the main allegations and claimed exposure:
| Issue | Allegation | Potential exposure |
|---|---|---|
| Copyrighted works | Tens of thousands of songs and lyrics allegedly used without permission | Up to $150,000 per work |
| Copyright data removal | Copyright management information allegedly stripped from works | Up to $25,000 per instance |
| Data sourcing | Alleged use of BitTorrent, pirate libraries, and lyric services | Could drive damages into billions |
| Defendants | Anthropic plus co-founders Dario Amodei and Benjamin Mann | Expanded liability risk |
How does this fit into the wider AI copyright fight?
This lawsuit fits into a much broader confrontation between generative AI companies and the industries that supply the material those systems are built on. Publishers, artists, record labels, and news organizations have been pressing courts to decide whether training on copyrighted works is fair use, infringement, or something in between.
So far, the industry has seen a mixed landscape. Some AI companies have struck licensing deals to reduce litigation risk, while others have fought aggressively in court. The legal uncertainty has encouraged a patchwork response: settlements in some areas, licensing in others, and high-stakes lawsuits where rights holders believe the training data was gathered unlawfully or used too extensively.
Anthropic’s situation is notable because the company has marketed itself as a safety-focused AI developer and has become one of the most prominent challengers to OpenAI in the foundation-model market. That profile makes it both commercially important and legally exposed. The larger the company’s footprint, the more incentive creators and publishers have to bring claims that could force licensing deals or major settlements.
The music industry also has a particular reason to pay close attention. Lyrics are short, recognizable, and commercially valuable, which makes them easier to identify in outputs and easier to argue about in court. If judges begin treating lyric training as straightforward infringement rather than a permitted form of model development, it could reshape how AI systems are built and licensed.
Why copyright management information matters
Copyright management information, or CMI, includes data such as the author’s name, the rights holder, and licensing terms attached to a work. If a company removes that information, the law can provide separate penalties because the action can undermine the system that tells users how content may be used.
In practical terms, CMI claims can be powerful because they do not depend solely on proving that a work was copied. A plaintiff may argue that removing identifying data made it harder to track ownership, prevented licensing, or concealed the source of protected material. That is one reason the publishers’ request for additional statutory damages could become significant.
How could the case affect Anthropic?
Anthropic could face substantial financial and reputational consequences if the lawsuit advances and survives early procedural challenges. The company is already absorbing the cost of related legal disputes, and a new case alleging both mass infringement and the removal of copyright data could make settlement pressure even stronger.
There is also a business implication beyond damages. Enterprise customers, investors, and partners may view repeated copyright disputes as a sign that the cost of using frontier AI remains uncertain. Even if Anthropic ultimately prevails on some claims, the legal defense burden alone can be expensive and distracting.
If the publishers succeed, the case could encourage more rights holders to demand upfront licensing for training datasets rather than pursuing damages after the fact. It could also strengthen the argument that AI firms need clearer records showing exactly what content was used, how it was obtained, and whether permission existed.
What happens next?
The immediate next step is Anthropic’s response in court. The company can be expected to challenge the allegations on both factual and legal grounds, potentially disputing the claims about how the material was collected, whether the works were actually used in model training, and whether the law permits such use.
Over time, the case could move through motions to dismiss, discovery, and possibly settlement talks. If it reaches trial, the court may be asked to decide not only whether infringement occurred, but also how to assess damages across a very large number of works and alleged acts of copying.
Because the complaint spans songs, lyrics, metadata, and alleged file-sharing practices, the litigation could become a broad test of AI training practices rather than a narrow dispute about a few isolated works. That makes it especially important for other model developers watching the case closely.
Timeline of the dispute and related actions
The music-industry fight against Anthropic did not begin with this complaint, and the broader legal pressure is still building. The following timeline shows how the conflict has escalated.
| Date/Period | Event | Why it matters |
|---|---|---|
| Earlier litigation | Anthropic faced a lawsuit from book publishers | Ended in a reported $1.5 billion settlement |
| Subsequent cases | UMG, Concord, ABKCO, BMG, and Round Hill Music brought separate actions | Shows industry-wide resistance to AI training practices |
| Latest filing | Sony Music and Warner Chappell sued in Northern California | Expands the dispute to another major rights-holding bloc |
Why the music business is drawing a harder line
The music business has long relied on licensing discipline to monetize recordings and compositions. Unlike some other creative sectors, it has deep experience with royalties, rights clearance, and negotiated use. That means publishers are especially likely to see unlicensed AI training as not just a legal violation but a direct threat to the licensing economy that supports songwriters and catalog owners.
There is also a cultural dimension. Songs are not generic text; they are memorable creative works with identifiable authors and commercial value. When a model can reproduce distinctive lyrics or closely echo a song’s structure, the industry may see that as more than abstract data use. It becomes, in the eyes of rights holders, a substitute for the original market.
That is why this suit may resonate beyond the courtroom. The question is no longer whether AI companies can access vast amounts of public material. The question is whether they must pay for it in the same way broadcasters, sample-based musicians, streaming services, and digital platforms have learned to do over time.
Key allegations at a glance
- Anthropic is accused of using tens of thousands of copyrighted music works without authorization.
- The suit names co-founders Dario Amodei and Benjamin Mann in addition to the company.
- Publishers claim the company relied on pirated books, torrents, and licensed lyric sources without proper permission.
- Specific songs cited include classics and contemporary hits from Marvin Gaye to Taylor Swift.
- The plaintiffs say statutory damages could total several billion dollars.
What this means for the AI industry
This case could become another benchmark in the evolving law of AI training. If courts increasingly accept the argument that model builders must license creative works before using them at scale, the economics of generative AI could shift substantially. Training datasets would become more expensive, more documented, and more heavily negotiated.
If, on the other hand, Anthropic successfully defeats the claims, the decision could reinforce the industry’s position that some kinds of large-scale data use are lawful, even when copyrighted content is involved. That outcome would likely embolden other AI companies to keep fighting similar cases and limit the scope of licensing demands.
For now, the lawsuit is another reminder that the AI boom is colliding with long-established creative industries that have little interest in giving away their catalogs for free. The legal outcome will not only affect Anthropic; it may help define the boundaries of how next-generation AI systems are built.
As this dispute unfolds, the central issue will remain whether training a generative model on protected creative works can be defended as innovation, or whether courts will treat it as the kind of large-scale copying that copyright law was designed to stop.
Frequently asked questions
Why did Sony Music and Warner Chappell sue Anthropic?
They sued Anthropic over allegations that the company copied tens of thousands of copyrighted songs and lyrics to train and operate its Claude AI models without permission. The publishers say the conduct was large-scale, commercially motivated and unlawful.
How much could Anthropic have to pay if the publishers win?
Anthropic could face billions of dollars in potential damages if a court awards the maximum statutory penalties sought by the plaintiffs. The complaint asks for up to $150,000 per work and additional penalties for alleged removal of copyright information.
Which Anthropic executives are named in the lawsuit?
The lawsuit names co-founders Dario Amodei and Benjamin Mann as individual defendants. The publishers say their involvement matters because the alleged infringement was not isolated but part of the company’s broader data collection and model-building strategy.
What songs are mentioned in the complaint?
The complaint cites several well-known songs, including 'Ain’t No Mountain High Enough,' 'Livin’ On a Prayer,' 'September,' 'Hallelujah' and 'Paper Rings.' The publishers say those works were found in Anthropic’s training data or related materials.
How does this lawsuit fit into the wider AI copyright battle?
It is part of a broader wave of copyright disputes over generative AI training data. Anthropic is already facing or has recently resolved other cases involving books and music, and the outcome could influence licensing and litigation across the AI industry.









