LLM Training Data and Fair Use: What the 2024–2025 Cases Left Unresolved

If you have been watching the AI copyright litigation landscape, you might think the courts have been busy settling things. And in one narrow sense, you would be right: by mid-2026, federal courts have issued a handful of substantive rulings on whether training large language models on copyrighted material constitutes copyright infringement. A $1.5 billion settlement in one major case made headlines. Two summary judgment wins for AI companies generated optimistic headlines about fair use.

But here is what those headlines missed: the courts that ruled in favor of AI companies were careful to say they were deciding the facts before them — not writing the rulebook for the industry. The judges who found fair use went out of their way to note that different facts might produce a different result. The most important questions in AI copyright law remain open, and in some cases, the litigation is just getting started.

For small and mid-sized businesses that are building on AI platforms, integrating AI tools into their workflows, or making decisions about training custom models, the legal uncertainty is not merely academic. It has direct consequences for vendor contracts, insurance coverage, and the exposure you may be absorbing without knowing it.

This post walks through the key cases, what they actually decided, and — more importantly — what they left unresolved.


The Cases So Far: A Tour of the Litigation Landscape

Thomson Reuters v. ROSS Intelligence — The First Fair Use Ruling

The first U.S. court decision to reach the fair use question in an AI training context came not from Silicon Valley but from the District of Delaware. On February 11, 2025, Judge Stephanos Bibas (sitting by designation from the Third Circuit) issued partial summary judgment for Thomson Reuters in Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc.

The facts were specific: ROSS Intelligence, a legal research AI startup, had used unauthorized copies of Westlaw’s editorial headnotes — short, lawyer-written summaries of judicial holdings — to train its competing AI legal research tool. The court found that ROSS had infringed 2,243 Westlaw headnotes and rejected ROSS’s fair use defense. The key factors that sank the fair use argument were (1) ROSS used the material to build a product that directly competed with Thomson Reuters’s own product, and (2) the use clearly harmed the existing and potential market for Thomson Reuters’s headnotes and derivative works.

This ruling is significant but also narrow: it involved a direct competitor using someone else’s content as a shortcut to build a substitute product. That is about as unfavorable a fact pattern as you can have on the fair use spectrum.

The case is now on appeal before the Third Circuit. Oral argument was scheduled for June 11, 2026. Among other issues, the Third Circuit will have to decide how much copyright protection attaches to short editorial summaries and how the market harm analysis should work when a startup uses training data to enter a new market rather than reproduce content for end users. This will be the first appellate opinion on AI training fair use from any U.S. circuit court, and it is worth watching closely.

Kadrey v. Meta — Fair Use Wins, But on Narrow Grounds

On June 25, 2025, Judge Vince Chhabria of the Northern District of California granted summary judgment in favor of Meta in Kadrey v. Meta (also known as Silverman v. Meta). The plaintiffs — a group of authors including Sarah Silverman, Richard Kadrey, and Christopher Golden — alleged that Meta used pirated copies of their books, sourced from shadow library sites such as Z-Library and Bibliotik, to train its LLaMA 1 and 2 language models.

Judge Chhabria’s ruling dismissed the case, but the grounds are important to understand. The decision was not a sweeping endorsement of AI training as fair use. Rather, the judge found that the plaintiffs had not presented meaningful evidence that any LLaMA output actually reproduced or was substantially similar to their copyrighted works. No output evidence, no infringement.

Crucially, Judge Chhabria explicitly left open the question of what would happen with a stronger evidentiary record. He commented that a plaintiff who could demonstrate market dilution or produce LLaMA outputs that replicated their work might well prevail. He did not hold that using pirated shadow library books to train AI models is categorically acceptable — he held that these plaintiffs had not connected the training to any provable harm.

That is a narrow holding dressed up in favorable language for AI companies.

Bartz v. Anthropic — The $1.5 Billion Settlement

The most consequential case to date did not end with a trial verdict. Bartz v. Anthropic (the consolidated author class action against Anthropic in the Northern District of California before Judge William Alsup) settled in August 2025 for $1.5 billion, covering approximately 482,000 works.

But before the settlement, Judge Alsup issued rulings that have shaped the entire subsequent litigation landscape. In June 2025, he held that Anthropic’s use of legally acquired books to train its Claude models was “quintessentially transformative” and likely constituted fair use. At the same time, he held that downloading and retaining pirated copies of books was not protected by fair use — it was standalone infringement, regardless of what the downloaded material was later used for.

That two-part analysis has become something of a template. It separates the training act from the acquisition act and treats them as legally distinct. An AI company that trains on pirated material faces a much harder case, even if the training itself might have been permissible had the material been obtained legitimately.

The implied per-book settlement rate in Bartz — roughly $3,100 per work — is also significant. It demonstrates that statutory damages exposure in AI training cases, even at negotiated settlement rates far below the $150,000 per-work statutory maximum for willful infringement, can reach numbers that threaten the financial viability of AI companies.

The OpenAI Consolidated Litigation — Output Claims Survive

In April 2025, the Judicial Panel on Multidistrict Litigation centralized multiple author copyright cases against OpenAI and Microsoft in the Southern District of New York as In re OpenAI, Inc. Copyright Infringement Litigation, MDL No. 3143. The Authors Guild case was folded into this proceeding.

On October 27, 2025, Judge Sidney Stein denied OpenAI’s motion to dismiss the consolidated class plaintiffs’ direct infringement claim based on ChatGPT’s outputs. The court found that the plaintiffs had plausibly alleged that ChatGPT’s outputs were substantially similar to their copyrighted works. This is not a finding of liability — it is a ruling at the pleading stage that the theory is legally viable and can proceed to discovery. But it is important because it establishes that output-based infringement claims are not going to be dismissed at the gate.

That case is now in active discovery, with no trial date set as of mid-2026.

The Music Industry Cases — Suno, Udio, and the Licensing Market in Real Time

In June 2024, the Recording Industry Association of America, on behalf of Universal Music Group, Sony Music, and Warner Music Group, filed separate lawsuits against Suno (in the District of Massachusetts) and Udio (in the Southern District of New York), alleging that both AI music generation companies trained their models on copyrighted sound recordings without authorization.

The litigation has fractured along settlement lines. Warner Music Group settled with Suno in November 2025. Universal Music Group settled with Udio in October 2025, establishing a licensing arrangement with per-generation royalty rates and content identification requirements. As of mid-2026, Sony Music has not settled with either company. Its cases remain active, with a potentially pivotal ruling on fair use expected from Judge Denise Casper in Massachusetts in summer 2026. If Suno wins on fair use, it could undermine the licensing arrangements that UMG and Warner have already struck. If Suno loses, those deals become the industry standard template.

Meanwhile, in January 2026, UMG, Concord, and ABKCO filed a separate $3 billion lawsuit against Anthropic — not over book training data, but over more than 20,000 copyrighted song lyrics. The complaint alleges that Anthropic engaged in mass piracy, downloading millions of unauthorized copies of works through BitTorrent and similar channels. Evidence produced in the Bartz case reportedly showed that Anthropic personnel personally authorized the use of pirate library sources. This case is being called the largest single non-class-action copyright case in U.S. history. It is in early stages as of mid-2026.


The Four Fair Use Factors and How Courts Have Applied Them

Under 17 U.S.C. § 107, courts evaluate fair use through four factors: (1) the purpose and character of the use, including whether it is commercial or nonprofit; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used; and (4) the effect on the potential market for or value of the copyrighted work.

In the AI training cases decided so far, the courts have struggled most with factors one and four.

Factor one — transformative use. Courts in Bartz and Kadrey found AI training to be highly transformative: the purpose of ingesting millions of books is not to reproduce books, but to teach a model statistical patterns in language. The model does not “remember” or “store” books as books; it encodes mathematical representations. By contrast, in Thomson Reuters v. ROSS, the use was non-transformative because ROSS’s output was a direct substitute for the original work — legal research answers competing with Thomson Reuters’s own legal research answers.

Factor two — nature of the work. Published, creative fiction gets strong copyright protection. Factual compilations get weaker protection. Most AI training datasets are heavily weighted toward published expressive works, which means this factor typically weighs against AI companies.

Factor three — amount used. Training on entire books is, by definition, copying 100% of each work. Courts in Bartz and Kadrey did not hold this to be disqualifying because the purpose was transformation, not reproduction. But it remains a factor that plaintiffs can exploit, particularly where the training data includes full copies obtained from pirate sources.

Factor four — market harm. This is the most contested and the most consequential factor going forward, for reasons addressed below.


Unresolved Question 1: Output Memorization as a Separate Infringement Theory

The most technically complex unresolved issue in AI copyright law is what happens when a language model reproduces training data verbatim in its outputs.

Both Judge Chhabria in Kadrey and Judge Alsup in Bartz went out of their way to note that the absence of output similarity evidence was central to their fair use rulings. The implication was direct: had the plaintiffs demonstrated that LLaMA or Claude outputs reproduced significant passages from their books, the analysis might have gone differently.

This is not a hypothetical concern. Academic research has consistently demonstrated that language models can be prompted — sometimes through seemingly innocuous queries — to reproduce verbatim text from their training data. A 2025 paper demonstrated methods for extracting long blocks of copyrighted book text from production-grade language models. The legal question this raises is whether output reproduction constitutes a second, independent act of infringement, separate from any question about whether the training itself was fair use.

The theory is that even if training was protected, each time a model outputs copyrighted text to a user, that output is an unauthorized reproduction or distribution of the original work. The user sees the original text; the AI is the delivery mechanism. Under 17 U.S.C. § 106, copyright owners hold exclusive rights not only over reproduction but also over public distribution and derivative works. Each verbatim output could be a separate act of infringement.

Judge Stein’s October 2025 ruling in the OpenAI case allowed this theory to survive a motion to dismiss. He found that the Authors Guild plaintiffs had plausibly alleged substantial similarity between ChatGPT outputs and their books. That case is now in discovery, which means the factual record on output similarity will be developed under oath.

In November 2025, the Munich I Regional Court in Germany actually found, in a GEMA case against OpenAI, that GPT-4’s reproduction of song lyrics in response to user queries constituted unlawful reproduction under German copyright law. That decision is under appeal, and European copyright law has different contours than U.S. law. But it illustrates that courts in multiple jurisdictions are taking output reproduction seriously as a distinct legal problem.

For U.S. law, the critical unresolved question is what standard of similarity triggers infringement in the output context. Copyright infringement requires more than trivial similarity — courts apply a “substantial similarity” test that asks whether an ordinary observer would recognize the similarity between the output and the original work. The challenge with AI outputs is that they can reproduce large passages exactly, produce close paraphrases, or generate content that is functionally equivalent without being verbatim. Different theories of substantial similarity could cover some or all of these scenarios.

Until an appellate court addresses this directly, every company deploying a large language model in a customer-facing application is operating with significant uncertainty about its output liability.


Unresolved Question 2: The Growing Licensing Market and Its Effect on Fair Use

The fourth fair use factor — market harm — is the one that could undo the favorable training-data rulings over time, and it is moving in real time.

Under copyright law, fair use is less available when a functioning market for licensing the work already exists. The Supreme Court articulated this in Campbell v. Acuff-Rose Music, 510 U.S. 569 (1994), and the principle has been applied consistently since: if the copyright owner can, in practice, license the use, courts are reluctant to say the use is free.

When the early AI training cases were filed in 2023, there was essentially no established licensing market for LLM training data in books or music. AI companies could credibly argue there was no market to harm. That argument is becoming harder to make.

In the music space, UMG’s settlement with Udio created an explicit licensing framework with per-generation royalty rates for generative AI trained on licensed catalog. Warner’s settlement with Suno includes a licensing component. UMG announced plans for a new subscription service in 2026 specifically for generative AI trained on fully authorized recordings. These are not theoretical licensing possibilities — they are actual market transactions at published terms.

In the books and text space, publishers have been building out similar frameworks. The Association of American Publishers and major houses have pushed for opt-in licensing rather than opt-out, which, if adopted as the industry standard, would create a robust licensing market for LLM training text.

The Copyright Office took note of this in its May 2025 report on AI training and fair use. That report, the third and final installment in the Office’s multi-year AI and copyright study, stated that uses for which licensing is reasonably available are less likely to qualify as fair use. The Office specifically identified commercial AI training on expressive works as an area where, in its view, fair use protection is weakest when licensing is available and the AI output competes in the market for the original work.

This is not binding on courts, but the Copyright Office’s views on copyright law carry significant weight with federal judges, who regularly consider Office guidance when interpreting statutory ambiguities.

The practical implication is significant: every licensing deal the music industry strikes with an AI company, and every book licensing framework that publishers construct, makes the fair use defense weaker for the next AI company that argues it had no obligation to license. The market is essentially being built out from underneath the fair use defense.

Sony Music’s pending case against Suno will be the first real test of this theory in the music space. If Judge Casper finds that Suno’s training was not fair use — in part because UMG and Warner had already established that licensing was commercially available — that ruling could reverse the entire narrative about AI training and fair use in the music industry.


Unresolved Question 3: Statutory Damages Exposure Even Without Final Liability

Even setting aside the question of whether training AI models ultimately constitutes infringement, the statutory damages structure under 17 U.S.C. § 504 creates enormous financial exposure for AI companies — and by extension, for businesses that use or deploy AI platforms that may themselves face liability.

Under current copyright law, a plaintiff who timely registered its copyright before infringement occurred can elect statutory damages rather than actual damages. Statutory damages range from $750 to $30,000 per work infringed, with no proof of actual harm required. For willful infringement — where the defendant knew or should have known its conduct was infringing — the ceiling rises to $150,000 per work.

Apply those numbers to a training dataset. The Anthropic book training dataset that gave rise to Bartz covered approximately 482,000 registered works. At the statutory maximum for willful infringement, the theoretical exposure would exceed $72 billion. That number is not realistic as a litigation outcome — courts have discretion to set amounts within the range — but it illustrates why defendants settle cases even when they believe they have strong fair use arguments. The downside risk of a trial loss is existential.

The UMG-Concord-ABKCO suit against Anthropic over song lyrics claims more than $3 billion based on over 20,000 songs. That is closer to the actual exposure range for a well-registered catalog.

Statutory damages exposure also creates a major asymmetry between large AI companies and the businesses that use their products. A company like Anthropic or OpenAI has resources to litigate and settle; a small startup building on top of their API does not. The question of whether downstream users of AI systems bear any copyright liability for their use of AI-generated outputs is not yet resolved by any court. But if courts eventually hold that AI outputs infringe, the question of who is liable — the AI developer, the deploying business, or both — will become critical.


What the Circuit Courts and Potentially the Supreme Court Still Need to Resolve

The current state of AI copyright law can be summarized as: a collection of district court decisions, none of which bind other district courts, with no circuit court opinions yet, and a Supreme Court that has so far declined to engage with AI-specific copyright questions (as it did in March 2026 when it denied certiorari in Thaler v. Perlmutter, the AI-generated authorship case).

Here is what appellate courts need to address before the law can stabilize:

The standard for transformativeness in AI training. The Bartz and Kadrey courts found AI training transformative; the ROSS court found it non-transformative. The difference turned largely on whether the AI output competed with the original work. The Third Circuit’s pending review of ROSS will be the first appellate opinion on this question, but it involves a non-generative AI system, so its application to LLMs may be limited.

The relationship between training fair use and output infringement. Courts have not yet squarely addressed whether a fair use finding on the training side protects the AI company from output-based infringement claims, or whether those are independent theories. This question will work its way up through the OpenAI MDL and other active cases.

The market harm analysis in the presence of an emerging licensing market. As licensing deals proliferate, courts will need to determine at what point a licensing market becomes sufficiently established to defeat a fair use defense. The Sony v. Suno ruling expected in summer 2026 may be the first decision to grapple with this directly.

The piracy distinction. Judges Alsup and Chhabria both suggested that acquiring training data through piracy is a standalone act of infringement, regardless of how the data is later used. This principle needs appellate validation and may eventually require the Supreme Court to define the boundaries between the acquisition question and the training question.

The circuit split that is coming. The ROSS case is in the Third Circuit. The major generative AI cases are in the Northern District of California (Ninth Circuit) and the Southern District of New York (Second Circuit). These circuits have meaningfully different copyright traditions, particularly on the transformativeness analysis. The Second Circuit’s approach to substantial similarity and the Ninth Circuit’s approach to fair use have historically diverged, and AI cases will eventually produce a circuit split requiring Supreme Court resolution.


Practical Implications for AI Businesses

If you are building on AI platforms, integrating AI tools into your products, or evaluating whether to train a custom model, here is what the unresolved legal landscape means for your business today.

Disclose Your AI Stack

Copyright liability in AI training is heavily fact-specific. Courts care deeply about where training data came from. If you are a business deploying a third-party AI system, understanding what data that system was trained on is not just a legal exercise — it is a business risk assessment. Many enterprise AI vendors now offer training data disclosures or documentation, but you may have to ask for it. Build this into your vendor evaluation process.

If you are training or fine-tuning a model yourself, document where your training data came from, how it was obtained, and what licensing or fair use analysis supports your use. That documentation will matter if you ever face a claim.

Vendor Contracts and Indemnification — Read the Fine Print

The indemnification provisions in AI vendor contracts have become one of the most consequential terms in enterprise software. Most standard vendor agreements cap the vendor’s liability at a fraction of annual fees and exclude intellectual property indemnification for content you input or that the AI generates in response to your prompts.

Ask these specific questions of any AI vendor you are evaluating:

  • Does the vendor provide intellectual property indemnification for copyright claims arising from the model’s training data?
  • What is the cap on the vendor’s indemnification obligation?
  • Does the vendor carry IP indemnification insurance to fund that obligation?

Some major vendors — Microsoft Copilot, for example — have announced that they will defend customers against copyright claims arising from the AI-generated outputs of their systems, subject to terms and conditions. Others have not made similar commitments. The gap between vendors on this issue is significant and consequential for your risk allocation.

Be especially cautious of the mismatch between vendor indemnification language and insurance coverage. A vendor’s agreement to indemnify you is only as good as the vendor’s ability to pay. Many AI vendors, particularly smaller ones, carry no insurance coverage specifically designed to cover third-party copyright claims in AI outputs. If the indemnifying vendor cannot fund its obligation, you may be left holding the liability.

The standard commercial general liability policy does not cover intellectual property infringement claims. Technology errors and omissions (Tech E&O) coverage, which is designed for technology vendors and suppliers, covers IP claims arising from the technology’s own services. If you are a customer deploying a vendor’s AI tool — rather than the vendor itself — your E&O policy may not respond to a copyright claim arising from AI outputs.

A growing market of specialty AI liability insurance is emerging in 2025 and 2026, designed specifically for businesses deploying AI systems that may generate infringing outputs. These products are new, coverage terms vary widely, and the claim history needed to price them accurately does not yet exist. But if your business relies on AI-generated content in any customer-facing application, the question of whether you have coverage for copyright claims arising from those outputs is worth exploring with your insurance broker.

Registered Works Versus Unregistered Works

If you are a content creator who believes your works may be in AI training datasets, timely copyright registration matters enormously. Only timely registered works are eligible for statutory damages and attorney’s fees. Without registration, a plaintiff is limited to proving actual damages — which, for most individual creators, is practically impossible to quantify in the AI training context. The registration requirement has become a critical filtering mechanism in these cases.

Watch the Sony v. Suno Ruling

For any business in the music technology, media, or content licensing space, the Sony Music case against Suno is the single most important pending development in AI copyright law. If Suno loses on fair use — in a case where UMG and Warner have already struck licensing deals, creating a functioning licensing market — it will establish that AI music generation companies must license training data. That precedent would extend to other AI applications that generate expressive content in markets where rights-holders have begun building licensing infrastructure.


Conclusion

The AI copyright cases of 2024 and 2025 gave us some answers and a great many more questions. We know that AI training is not automatically fair use and not automatically infringement — it depends heavily on the facts, including where the data came from, whether the output competes with the original work, and whether a licensing market exists. We know that downloading training data through piracy is not protected. We know that output-based infringement claims can survive a motion to dismiss. We know that statutory damages exposure is large enough to threaten even well-funded AI companies.

What we do not yet know — and what the courts have not yet resolved — is where the fair use line falls in the presence of a growing licensing market, whether a fair use finding on training immunizes a company from output infringement claims, and how circuit courts will apply their own established copyright doctrines to the AI training question. Those answers will come from the Third Circuit appeal of ROSS, from Judge Casper’s ruling in the Suno case, from the OpenAI MDL in New York, and eventually, likely, from the Supreme Court.

For businesses in the meantime, operating in this uncertainty requires discipline: document your data sourcing practices, scrutinize the indemnification terms in every AI vendor agreement, evaluate your insurance coverage for IP claims, and watch the developing case law closely. The legal framework is being built in real time, and the decisions coming in the next 18 months will define it for the decade ahead.


This post is for informational purposes only and does not constitute legal advice. If you have questions about your organization’s specific AI copyright exposure, please consult qualified legal counsel.



Leave a Reply