AI Training Data and Copyright: What the 2024-2025 Cases Settled and What They Left Open
- August 9, 2026
- Posted by: allan
- Category: Uncategorized
Every company that builds or deploys AI products is now operating in contested legal territory. The question of whether it is lawful to train a machine learning model on copyrighted text, code, images, or recordings has moved from academic debate to federal courtrooms. Between February and September 2025, a series of rulings produced the first real body of American caselaw on this issue — and the results are anything but uniform. Two courts reached opposite conclusions about fair use in a span of four months. The music industry settled some cases while others remain in active litigation. And the single biggest case in the field — the New York Times against OpenAI — has not yet produced a merits ruling at all.
For business owners and startup founders deploying AI or building AI-powered products, these cases are not abstract. They define the legal risk profile of your technology stack and your training pipeline. This post walks through what the 2024–2025 cases actually held, what they left open, and what practical steps your organization should be taking now.
Why This Issue Matters for Every AI Company
When a language model, an image generator, or a music-creation tool is trained, it ingests enormous quantities of existing content. The training process involves copying that content — typically at scale — and extracting statistical relationships from it. Whether that copying constitutes copyright infringement depends primarily on whether it qualifies as fair use under 17 U.S.C. § 107.
Fair use is not a loophole. It is a doctrine courts have applied for over a century to allow uses of copyrighted material that serve the public interest without requiring permission — scholarship, commentary, parody, and transformative creative work. The question that has consumed AI litigation is whether ingesting millions of books, articles, images, or recordings to build a commercial AI product fits within that doctrine.
If it does, AI developers need no licenses. If it does not, virtually every major AI model in existence was built on infringement, and the owners of those rights have potential claims worth trillions of dollars in the aggregate. The stakes are enormous, which explains why the litigation has attracted some of the most prominent plaintiffs and the largest law firms in the country.
The Fair Use Doctrine and the Four-Factor Test
Fair use under the Copyright Act requires courts to weigh four statutory factors. No single factor is automatically dispositive, though courts often describe the fourth — market harm — as the most important. The four factors are:
Factor 1: Purpose and Character of the Use. The central question is whether the use is “transformative” — does it add new expression, meaning, or message, or does it merely supersede the original work? Commercial uses are presumptively less favored, though they are not automatically unfair. Courts have also distinguished between using a work to learn (which may be transformative) and using a work to produce a substitute (which usually is not).
Factor 2: Nature of the Copyrighted Work. Fair use is more readily available for factual or informational works than for highly creative ones. A training set composed primarily of news articles sits differently than one composed of literary novels or original musical recordings. Where the copyrighted works are the product of significant creative effort — fiction, poetry, original composition — this factor tends to weigh against the alleged infringer.
Factor 3: Amount and Substantiality of the Portion Used. Training at scale typically involves copying entire works, not excerpts. Wholesale copying weighs against fair use, though it is not conclusive. Courts have recognized that sometimes an entire work must be copied for a purpose to be achieved — a point that cuts both ways in the AI training context.
Factor 4: Effect on the Potential Market. Courts have called this the single most important factor. The relevant question is not just whether the use harms sales of the original work, but whether it harms the market for potential derivative works — including licensed uses. If a market exists (or could plausibly exist) in which copyright owners could license their works for AI training, and a defendant’s unlicensed use displaces that market, this factor weighs heavily against fair use.
These four factors have been applied inconsistently across the 2024–2025 AI training cases, producing a landscape that is genuinely uncertain.
Thomson Reuters v. Ross Intelligence: The First Federal Ruling
What the Case Was About
The litigation between Thomson Reuters — owner of the Westlaw legal research platform — and Ross Intelligence began in 2020. Ross had built an AI-powered legal research tool and, after being denied a license to use Westlaw content directly, worked through a third-party vendor to obtain training data derived from Westlaw’s proprietary headnotes. Headnotes are the brief editorial summaries that appear at the top of case law on Westlaw; they are written by Thomson Reuters attorneys and represent a significant creative and economic investment.
On February 11, 2025, the United States District Court for the District of Delaware granted summary judgment in favor of Thomson Reuters on the copyright claims. Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., No. 1:20-cv-613-SB (D. Del. Feb. 11, 2025). This was the first federal ruling to squarely decide whether training an AI system on copyrighted material is or is not fair use.
The Copyrightability Finding
Before reaching fair use, the court addressed whether the headnotes were even protectable. The threshold for copyright protection is low — there need only be a minimal degree of creativity. The court held that the headnotes met that standard. They are not mere reproductions of judicial language; they involve editorial synthesis, distillation, and explanation. The court found that 2,243 out of 2,830 headnotes at issue had been substantially copied.
The Fair Use Analysis
The court analyzed all four factors, but its ruling turned almost entirely on the fourth.
On the first factor, the court found that Ross’s use was not transformative. Ross was building a legal research product — the same market that Westlaw serves. The fact that a machine learning model rather than a human was consuming the headnotes did not change the commercial purpose of the end product. The court rejected the argument that training a model is inherently transformative because the model output is different from the training input.
On the second factor, the court found the headnotes were sufficiently creative to weigh against Ross, given the editorial judgment involved in writing them.
On the third factor, the court noted that the copying was substantial.
On the fourth factor — and this is the holding that will shape AI copyright litigation for years — the court found that even though Thomson Reuters had not yet built or licensed its own AI training product using these headnotes, that did not matter. The court reasoned that a potential market existed for licensing Westlaw content for AI training purposes. Ross’s unlicensed use foreclosed that market. The court explicitly stated: “It does not matter whether Thomson Reuters has used the data to train its own legal search tools; the effect on a potential market for AI training data is enough.”
Why This Ruling Matters
The Ross decision is significant for several reasons. First, it is the first ruling to reject a fair use defense in the AI training context, establishing that such a defense is not automatic. Second, its market-substitution analysis is broad — it does not require the copyright owner to have already entered the AI licensing market, only that such a market could plausibly exist. Third, the court limited its ruling to non-generative AI, expressly noting that Ross’s tool was a search product, not a generative model that produces new text. Whether the same analysis applies to large language models that generate original content was explicitly left open.
The case is on appeal to the Third Circuit as of mid-2025.
The Authors’ Cases: Kadrey, Bartz, and the Ongoing Battle Over Books
Kadrey v. Meta
Filed in the Northern District of California in 2023, Kadrey v. Meta Platforms was one of the first and most prominent cases alleging that Meta had infringed copyrights by training its LLaMA family of large language models on books without authorization — including books obtained from “shadow libraries” like Library Genesis, which host pirated copies of copyrighted works.
In June 2025, the court granted summary judgment in Meta’s favor on the fair use question. The court found that Meta’s use was highly transformative — ingesting text to learn statistical relationships across language is a fundamentally different purpose from reading or distributing those books. The court also found that the plaintiffs had presented insufficient evidence of market harm. Critically, the court left a door open: it noted that future plaintiffs with “better-developed records on the market effects” of AI training might succeed where the Kadrey plaintiffs had not. The absence of proof was what sank the case, not the absence of a legal theory.
Bartz v. Anthropic
The Bartz v. Anthropic case, also in the Northern District of California, produced the most detailed fair use analysis yet in the AI training context. In June 2025, Judge William Alsup issued a split ruling that illustrated exactly how nuanced these cases are.
On the central question of whether training Anthropic’s Claude models on copyrighted books was fair use, the court ruled in Anthropic’s favor. The court found the use “exceedingly transformative” — Anthropic was not trying to reproduce or distribute the books; it was trying to build a general-purpose reasoning and language system. The court drew an analogy to a student who reads widely to develop knowledge and capability: authors cannot copyright the knowledge that readers extract from their works.
But the court drew a sharp line at piracy. Anthropic had obtained a portion of its training data by downloading books from pirate sites including Library Genesis and the Pirate Library Mirror — more than seven million books in total. The court strongly indicated that this conduct was not protected as fair use. Downloading pirated copies, the court reasoned, is not the same as digitizing lawfully acquired works; the manner of acquisition matters.
Despite winning on the fair use question for legitimate training data, Anthropic ultimately settled the case in August 2025. After Judge Alsup certified the case as a class action — bringing hundreds of thousands of copyrighted works within its scope and creating theoretical statutory damages liability exceeding $70 billion — the parties reached a class-wide settlement. The terms have not been fully disclosed as of this writing.
The Silverman v. OpenAI Track
Sarah Silverman and other authors filed against OpenAI in the Northern District of California in 2023. That case was partially dismissed at the pleadings stage and later consolidated with related author cases. In April 2025, the consolidated proceedings were transferred to the Southern District of New York, where they are now proceeding alongside related cases. The litigation is in active discovery — counsel for the authors have been permitted to inspect OpenAI’s training data in a secure facility. No summary judgment ruling on the fair use question is expected before 2026. This case remains one to watch precisely because it targets OpenAI’s most commercially prominent models and the broadest range of claimed damages.
The Music Industry: Settlements and Continuing Battles
The Recording Industry Association of America, on behalf of Universal Music Group, Sony Music, and Warner Music, filed suit in June 2024 against Suno and Udio — two AI music generation startups — alleging that both companies trained their models on copyrighted recordings without authorization.
The music cases present a particularly stark market-substitution argument. Unlike language models that generate text for general reasoning purposes, music generation tools produce outputs — songs — that directly compete with the works from which the training data was drawn. A user who prompts Suno to generate a pop song in a particular style is, in at least some economic sense, substituting AI-generated content for content they might otherwise license or purchase from a label.
The litigation proceeded on divergent tracks. Udio settled with the labels in late 2025, and Warner settled with Suno around the same time. Sony Music, however, continued to litigate against both companies through at least mid-2026. The settlements in the music cases produced licensing frameworks — in effect, the labels extracted royalty arrangements from the AI companies as a condition of settlement, creating a de facto licensing market in real time.
Independent artists also organized their own class actions. Country artist Tony Justice and his label filed suit in June 2025 against both Suno and Udio on behalf of independent artists and producers whose works appeared on streaming platforms since 2021.
The music litigation demonstrates a pattern that may recur across other content categories: major rights-holders use litigation as leverage to extract licensing deals, while leaving the underlying legal questions unresolved through settlement rather than merits rulings.
The New York Times Case: The Biggest Case With No Merits Ruling Yet
The lawsuit filed by The New York Times against OpenAI and Microsoft in December 2023 is the highest-profile AI copyright case in the country, but as of mid-2026 it has produced no ruling on the central fair use question. In March 2025, the Southern District of New York denied OpenAI’s motion to dismiss, allowing the copyright infringement claims to proceed. The case is now deep in discovery, with significant battles over OpenAI’s output logs and user data.
The Times’s theory is distinctive. In addition to the training data claim, the Times presented examples of ChatGPT reproducing Times articles nearly verbatim in response to user prompts — an output-side harm that goes beyond the training-side copying at issue in most other cases. Whether this output-level reproduction strengthens or complicates the fair use analysis is one of the questions that will ultimately need to be answered on the merits.
The Market Substitution Theory: The Most Consequential Question in AI Copyright Law
Running through all of these cases — Ross, Kadrey, Bartz, the music cases, the Times case — is a single underlying legal question that has not yet been definitively resolved: when, if ever, does training an AI on copyrighted content harm the “market” for that content in a way that defeats fair use?
The Ross court said a potential licensing market is enough. You do not need to show that the copyright owner was already selling licenses; you just need to show that the unauthorized use foreclosed a market that the owner could reasonably exploit. Under that logic, almost any AI training on commercial content is a potential problem, because content owners can always argue they should have been paid for the training use.
The Kadrey and Bartz courts cut the other way — if you cannot show actual market harm, you lose. The mere possibility that a licensing market could exist is not enough; you need evidence.
This tension is not yet resolved at the appellate level. The Third Circuit will have something to say about Ross. The Second Circuit will eventually hear appeals from the New York Times case and the Silverman/OpenAI consolidated proceedings. The Ninth Circuit may weigh in on the California cases. Until circuit courts — and potentially the Supreme Court — address this question, the doctrine will remain genuinely uncertain.
What is clear is that the fourth fair use factor is where AI training copyright cases are won and lost. Any company assessing its legal exposure should be building a record on this question: what is the market for the content you used? Did a licensing market exist when you trained? Did you displace it?
What These Cases Left Open
Despite the volume of litigation in 2024 and 2025, a remarkable number of important questions remain unresolved.
Generative AI and the Ross holding. The Ross court expressly limited its ruling to non-generative AI. Whether the same fair use analysis applies to a large language model that generates original text — rather than merely searching existing text — is an open question. The Bartz and Kadrey courts came out differently than Ross, but they were analyzing different facts in a different circuit. There is no consensus rule.
The piracy problem. The Bartz court signaled that downloading pirated training data is not protected conduct even if the same training on legitimately acquired copies might be fair use. How far this principle extends — and whether it taints an entire model trained on a mix of legitimate and pirated data — has not been decided.
Output-side infringement. Cases like the Times lawsuit raise the question of whether AI outputs that closely reproduce copyrighted training data are independently infringing, separate from the training process itself. This is analytically distinct from the training-data question and may produce a body of law that imposes liability even where the training itself is found to be fair use.
International variation. The cases discussed here are all American. Other jurisdictions — particularly the European Union and the United Kingdom — have different statutory frameworks for text and data mining. A company training a model in multiple jurisdictions cannot assume that a favorable US ruling protects it everywhere.
Works made for hire and corporate authorship. Most AI training data litigation to date has involved individual authors asserting rights in their books. The copyright landscape for training on code repositories, corporate websites, and other commercially generated content involves different ownership structures that have barely been litigated.
Practical Implications for Businesses
If You Are Deploying AI Built by Someone Else
When you use a third-party AI product — a language model API, an image generation service, a code assistant — you are not typically the entity that did the training. The copyright liability for the training data rests with the model developer, not the downstream deployer. However, your vendor contracts matter. Make sure your AI vendors represent that they have cleared rights in their training data or that they will indemnify you if training-data claims arise. Review the indemnification provisions in your API agreements carefully; many standard SaaS agreements exclude IP infringement claims relating to third-party training data or cap indemnification at low dollar amounts.
Output-side liability is a separate question. If your product causes an AI to reproduce copyrighted content verbatim in outputs that you distribute to users, that exposure may fall on you regardless of whether the training was lawful. Implement filtering or review mechanisms for outputs that might reproduce substantial portions of third-party content.
If You Are Building or Fine-Tuning Your Own Models
The cases make clear that the source of your training data matters enormously. Training on data obtained from pirate libraries or scraped without authorization from sites that prohibit scraping in their terms of service creates exposure that a fair use defense may not cure. The Bartz court drew that line explicitly, and other courts are likely to follow.
Before you train, conduct a training data audit. Identify the sources of your training corpus, assess whether those sources impose license restrictions, and document your analysis. If you are using a third-party dataset, understand how that dataset was compiled and what rights were cleared. The cases demonstrate that courts will look at the manner of acquisition, not just the purpose of the training.
Consider what market you are entering. The Ross case turned heavily on the fact that Ross was building a product that competed directly with the Westlaw platform from which it derived its training data. If your AI product competes with the industry that produced your training data, your fair use argument is substantially weaker than if your product serves an entirely different market.
Licensing Considerations
The litigation of 2024–2025 has produced an accelerating market for licensed AI training data. Getty Images and Shutterstock have established formal AI licensing programs. Publishers are negotiating licenses with model developers. The music industry is embedding royalty frameworks into its litigation settlements. The question is no longer whether licensed AI training data markets will exist; they exist now. The question is whether the cost of licensing is worth the legal certainty it provides.
For many companies, the answer is yes. The alternative — a training pipeline with contested provenance — creates the kind of open-ended legal exposure that can complicate fundraising, M&A due diligence, and enterprise sales. A documented licensing strategy is a significant competitive differentiator as AI procurement processes in large organizations increasingly require vendors to demonstrate legal compliance in their AI development practices.
The Emerging Market for Licensed AI Training Data
One of the most consequential practical developments of the 2024–2025 litigation wave has been the acceleration of formal AI training data markets. What began as ad hoc licensing negotiations between model developers and publishers has become a structured commercial market.
Shutterstock reported over $100 million in AI data licensing revenue from a single deal with OpenAI, with that deal potentially worth as much as $250 million through 2027. Getty Images launched a dedicated AI data licensing program and, following its announced merger with Shutterstock, now commands one of the largest licensed image libraries available for AI training. Associated Press, Reuters, and major book publishers have negotiated training data licenses with various AI developers. In the music industry, the post-litigation licensing frameworks negotiated with Suno and Udio established the principle that per-generation royalties are a viable commercial model.
This market development is not merely a business story. It has legal significance. Every licensing deal that gets signed strengthens the argument that a derivative market for AI training data exists — and thus strengthens the fourth-factor argument that unlicensed training displaces that market. As more rights-holders license their content, the harder it becomes for AI developers to argue that no licensing market exists for their training data.
The feedback loop between litigation and licensing will continue. Expect rights-holders to use the unresolved legal uncertainty as leverage in licensing negotiations, and expect AI developers to use licensing deals as evidence that they respect intellectual property — a posture that affects both litigation optics and regulatory relationships.
Conclusion
The 2024–2025 AI training data copyright cases produced the first real legal answers to questions that have lurked since large-scale machine learning became commercially significant. The Ross court said that training a non-generative AI on a competitor’s copyrighted content — where a licensing market plausibly exists — is not fair use. The Bartz and Kadrey courts said that training a generative AI on legitimately acquired books can be fair use, but training on pirated content is a different story, and proof of market harm matters. The music industry cases produced licensing deals that normalized royalty payments for AI training while leaving the underlying fair use questions unresolved.
What remains open is substantial: how appellate courts will reconcile the Ross holding with the California decisions, whether the New York Times output-side theory will succeed, how courts will handle models trained on mixed legitimate-and-pirated data, and whether Congress will ultimately legislate a compulsory licensing regime or some other statutory solution.
For business owners, the practical message is this: the legal landscape is still forming, but the direction of risk is clear. Provenance matters. How you obtained your training data, whether licensing markets exist for it, and whether your AI product competes with the source of its training data are the three questions that will determine your fair use exposure if litigation arises. Get ahead of those questions now — through training data audits, vendor contract review, and engagement with the emerging licensed data markets — rather than after a litigation demand arrives.
This post is intended for general informational purposes and does not constitute legal advice. If you have specific questions about AI copyright compliance, training data licensing, or vendor contract review, contact a qualified attorney.
