Generative AI's Collision with Copyright Law

Pamela Samuelson

International Conference on Machine Learning 2025 · Invited Talk

Overview

Pamela Samuelson, a distinguished legal scholar, delivered a compelling talk at ICML 2025, dissecting the intricate and often contentious relationship between generative AI technologies and established copyright law. The presentation illuminated the current legal landscape, marked by a surge of lawsuits challenging the legality of using copyrighted works as training data for large-scale AI models. Samuelson's core message underscored that copyright law is an undeniable and increasingly critical regulatory force for machine learning, a reality many in the technical community might wish to ignore but cannot.

Watch on SlidesLive

Visual summary for Generative AI's Collision with Copyright Law by Pamela Samuelson
Visual summary for Generative AI's Collision with Copyright Law by Pamela Samuelson

Key moments

  1. 0:00 Copyright law now regulating machine learning and AI.
  2. 2:00 Exclusive rights, fair use, and what copyright protects.
  3. 3:00 Fair use: balancing mechanism for technological change.
  4. 4:00 Copyright doesn't protect facts, ideas, or methods.
  5. 4:40 The four factors of fair use analysis.
  6. 6:00 Authors Guild v. Google Books fair use precedent.

Generative AI's Collision with Copyright Law

Speakers: Pamela Samuelson

Conference: ICML 2025

YouTube: https://slideslive.com/39043346

Overview

Pamela Samuelson, a distinguished legal scholar, delivered a compelling talk at ICML 2025, dissecting the intricate and often contentious relationship between generative AI technologies and established copyright law. The presentation illuminated the current legal landscape, marked by a surge of lawsuits challenging the legality of using copyrighted works as training data for large-scale AI models. Samuelson's core message underscored that copyright law is an undeniable and increasingly critical regulatory force for machine learning, a reality many in the technical community might wish to ignore but cannot.

The talk systematically addressed the fundamental principles of US copyright law, particularly the fair use doctrine, and then applied these principles to the ongoing legal battles involving major generative AI developers. Samuelson highlighted the profound financial stakes, with potential damages reaching billions or even trillions of dollars, and the significant uncertainty these cases introduce for the future of AI innovation. Her analysis provided crucial context on why creators—artists, authors, and programmers—are filing these lawsuits, driven by a blend of anger, a desire for fair compensation, and fear of market displacement.

This discussion is paramount for the AI/ML community, as the outcomes of these legal challenges will shape how models are trained, what data can be used, and the very business models of AI companies. Samuelson emphasized the need for technical professionals to understand these legal complexities and actively engage in policy debates, rather than passively observing. The talk served as a critical bridge between the technical advancements celebrated at conferences like ICML and the complex legal frameworks that govern their deployment and impact on society.

Background

▶ Watch: Copyright law now regulating machine learning and AI. (0:00)

The foundation of US copyright law, as enshrined in Article 1 Section 8 Clause 8 of the US Constitution, grants Congress the power to "promote the progress of science and useful arts" by securing exclusive rights to authors for limited times. This clause is interpreted as primarily serving the public good by fostering the creation and dissemination of knowledge, with the reward to authors being a secondary, albeit important, consideration. Copyright protects the original expression in works of authorship, such as books, art, and music, but explicitly not ideas, facts, methods, or styles. Authors are granted exclusive rights, most notably the right to control reproductions of their work.

However, these exclusive rights are not absolute. The fair use doctrine serves as a crucial limitation, providing "breathing space" for follow-on creations and allowing copyright law to adapt to unforeseen technological advancements. Fair use is not an excused infringement but rather a determination that a particular use is simply not an infringement at all. Courts evaluate fair use based on four key factors:

  1. Purpose and character of the use: Is it transformative? (i.e., does it add new meaning or purpose, or merely supersede the original?) Commerciality is also considered, though less weighted if the use is highly transformative.
  2. Nature of the copyrighted work: Factual works generally have "thinner" copyright protection than highly creative works.
  3. Amount and substantiality of the portion used: How much of the original work was copied, and was it the "heart" of the work? This is balanced against the purpose of the secondary use.
  4. Effect of the use upon the potential market for or value of the copyrighted work: Does the secondary use supplant demand for the original or for reasonable licensing markets?

A pivotal precedent in this area is Authors Guild vs. Google, where Google's digitization of approximately 20 million books for indexing and snippet display was deemed fair use by the Second Circuit Court of Appeals. The court found Google's purpose transformative (indexing for search, not expressive exploitation), the amount copied reasonable given the indexing purpose (entire books were needed), and no significant market harm because indexing did not supplant demand for the original books or a legitimate licensing market for such a use. Snippets were considered too short and scattered to undermine author incentives.

Internationally, the legal landscape varies. Israel, Japan, and Singapore have proactively enacted broad exceptions to copyright rules to support their burgeoning AI industries, often explicitly allowing text and data mining (TDM) for training purposes. The European Union adopted TDM exceptions a few years ago, distinguishing between non-profit research institutions (who have an opt-out-proof privilege) and other researchers (where copyright owners can opt out). However, the EU still faces challenges in establishing practical collective licensing schemes for global datasets. This patchwork of national laws highlights the global complexity of AI copyright.

Key Findings

▶ Watch: Fair use: balancing mechanism for technological change. (3:00)

The talk revealed a legal landscape in significant flux, with profound implications for the generative AI industry. Currently, over 44 lawsuits are pending in the US alone, challenging the use of in-copyright works for AI training, with similar cases in Canada and the UK. The most common claim is infringement of the reproduction right. Many of these are class action lawsuits, where lead plaintiffs claim to represent vast numbers of similarly impacted creators, aiming for potentially enormous damage awards. Prominent non-class action cases include Getty Images vs. Stability AI (challenging use of Getty's photography) and The New York Times vs. Microsoft and OpenAI (concerning news content). More recently, Disney sued Midjourney over outputs depicting movie characters.

The motivations behind these lawsuits are multifaceted:

  • Class action lawyers: Driven by the potential for massive contingency fees (e.g., a third of a potential $9 billion settlement mentioned in one case).
  • Large copyright owners: Primarily seek monetary compensation and greater control over their intellectual property. Many of these cases may eventually settle.
  • Individual creators (authors, artists, songwriters, programmers): Express profound anger and a sense of injustice that large corporations are profiting immensely from their work without permission or payment. They fear market competition from AI outputs that could undermine their livelihoods.

Crucially, Samuelson highlighted that the legal uncertainty in the US is likely to persist for several years, possibly a decade, due to the novel nature of these issues and the appellate process. The outcomes in US courts are expected to have significant ramifications globally, influencing how other countries decide to regulate AI training data.

Initial judicial decisions in two recent summary judgment cases—Bartz vs. Anthropic and Kadrey vs. Meta—offered some clarity but also stark divergences:

  • Common Ground: Both judges deemed the use of copyrighted works as training data "highly transformative" due to the fundamentally different purpose (extracting data/parameters vs. expressive consumption). Both also rejected the plaintiffs' argument that authors are entitled to control or license a market specifically for AI training data.
  • Key Divergences:
  • Pirated Books: Judge Alsop in Bartz found downloading pirated books for a training database not fair use, while Judge Chhabria in Kadrey considered the source of pirated books neutral to the fair use analysis for training purposes.
  • Market Dilution Theory: Judge Alsop dismissed the idea of "market dilution" (AI outputs flooding the market and undermining human creators' income) as "science fiction." In stark contrast, Judge Chhabria in Kadrey introduced it as a potentially "cognizable indirect harm," even though no precedent exists for it. He signaled that future plaintiffs should provide evidence for this theory.

These initial rulings, while influential, are far from final and are likely to be challenged on appeal, further prolonging the legal uncertainty.

Technical Deep Dive

▶ Watch: Copyright doesn't protect facts, ideas, or methods. (4:00)

While this talk is fundamentally about legal interpretation rather than machine learning architecture, the "technical deep dive" here focuses on the intricate legal mechanics and interpretations directly impacting AI systems and their development. The core technical "act" under scrutiny is the ingestion of massive datasets—often billions of copyrighted works—to train large language models (LLMs) and other generative AI systems.

The legal arguments revolve around how this ingestion process interacts with the exclusive right of reproduction and the fair use doctrine. Defendants, typically generative AI companies, argue their use is highly transformative because they are not exploiting the expressive purpose of the original works. Instead, they are using the works as "data" to extract patterns, relationships, and parameters that enable the model to generate new content. This contrasts sharply with traditional copyright infringement, which typically involves making substantially similar copies that directly compete with the original.

Key arguments from the defendants' perspective, drawing parallels to previous fair use cases:

  • Non-Expressive Purpose: AI models are designed to learn underlying statistical relationships and structures, not to "read" or "enjoy" the content in an expressive sense. This aligns with the transformative purpose found in Authors Guild vs. Google Books (indexing) and Field vs. Google (indexing and caching internet content).
  • Intermediate Copying: Companies argue that the copies made during training are intermediate steps necessary to build a non-infringing end product (the model). This draws on cases like Sega vs. Accolade, where intermediate copying of code for reverse engineering was deemed fair use if the final product was non-infringing. The model itself, they argue, does not contain "substantially similar expression" to any single training work.
  • Market Failure for Licensing: AI developers contend that obtaining licenses for billions of diverse copyrighted works from countless individual creators is practically impossible. This "market failure" argument suggests that if a rational licensing market cannot be formed, the socially valuable use should be permitted under fair use.
  • Output Filters: Some defendants, like Meta in the Kadrey case, highlighted the use of output filters as a technical safeguard to prevent the model from generating infringing content or "regurgitating" training data. This demonstrates an attempt to mitigate potential market harm.

Conversely, plaintiffs argue that AI training constitutes multiple exact reproductions of their entire works, often without permission, for commercial gain. They contend that the models do learn and can reproduce elements that compete with their work, even if not direct copies. The "nature of the copyrighted work" (often highly creative) and the "amount and substantiality" (entire works copied) are factors they believe weigh against fair use.

The emergent "market dilution" theory by Judge Chhabria introduces a new legal concept that directly challenges the technical outputs of generative AI. This theory posits that even if AI outputs are not substantially similar to specific input works, a flood of AI-generated content could indirectly collapse the market for human-created works, thereby constituting a cognizable harm. This theoretical harm, if proven, would require AI companies to demonstrate that their models do not contribute to such market collapse, potentially necessitating deeper scrutiny of model behavior and economic impact. The technical community would then face the challenge of providing empirical data on market effects, a task currently deemed difficult and speculative.

Experimental Setup & Results

▶ Watch: The four factors of fair use analysis. (4:40)

In the context of this legal talk, the "experimental setup" refers to the various lawsuits filed against generative AI companies, and the "results" are the initial judicial rulings and their reasoning. These legal "experiments" test the boundaries of copyright law, specifically the fair use doctrine, against the novel challenges posed by AI training and output generation.

Plaintiff Arguments (Hypotheses of Infringement):

  • Commerciality & Exact Copies: Plaintiffs argue that AI companies are commercial entities making exact, multiple copies of their works, often derived from "pirated books and other kind of infringing ill-gotten material," thus demonstrating "bad faith."
  • Nature of Work: Their works are highly creative, and AI developers chose them for their expressive value, not just as raw data.
  • Amount Copied: Copying entire works multiple times should weigh heavily against fair use.
  • Market Harm: Plaintiffs claim lost licensing revenues, lost sales, and unfair competition from AI outputs that either mimic or directly substitute their creative works.

Defendant Arguments (Fair Use Defense):

  • Transformative Purpose: The core defense is that AI training is "highly transformative," using works for a fundamentally different, non-expressive purpose (data extraction for model building). Commerciality is less relevant when the use is transformative.
  • Data, Not Expression: Defendants claim they don't care about the expression but merely use the works as data. Many works were publicly available via scraping, which is generally lawful.
  • Reasonable Copying: Copying entire works is "reasonable in light of the purpose" (e.g., to fully understand content for training).
  • No Substantial Similarity in Outputs: For the most part, AI outputs do not contain "substantially similar expression" to individual training works. They learn from the data but don't regurgitate it (and use filters to prevent this).
  • Infeasible Licensing: Licensing "billions of copyrighted works" from individuals is impossible, indicating a "market failure" that should permit socially valuable uses.

Key Case Results (Initial Judicial Rulings):

  1. Bartz vs. Anthropic (Judge Alsop):
  • Training Data Use: Found to be fair use. The purpose was "quintessentially transformative," and while books are creative, copying entire works was deemed "necessary" for the transformative purpose. No market harm was found, and authors were not entitled to control a licensing market for training data.
  • Digitizing Purchased Books: Also fair use, considered transformative, with reasonable copying and no market effect.
  • Downloading Pirated Books: Not fair use. The judge drew a clear line against using unlawfully obtained material.
  1. Kadrey vs. Meta (Judge Chhabria):
  • Training Data Use: Found to be fair use. Similar to Bartz, the use was "highly transformative," and commerciality was less weighted. The use of pirated books for training was considered neutral to the fair use analysis (contrasting Alsop).
  • Market Effects & "Market Dilution": This judge found no evidence of lost sales or licensing revenue from Kadrey. However, he introduced a novel concept: "market dilution," theorizing that AI-generated content could "flood the market" and cause it to "collapse" for human authors. While he found no evidence presented by Kadrey for this, he signaled that this could be a "cognizable" harm in future cases, effectively placing the burden on future plaintiffs to provide empirical data.

Precedent Cases Cited by Defendants:

  • Authors Guild vs. Google: Digitizing books for indexing and snippets was fair use (transformative purpose, no market harm).
  • Field vs. Google: Copying internet content for indexing and caching was fair use.
  • Plagiarism Detection Software Cases: Making and storing copies of student papers for plagiarism detection was fair use.
  • Sega vs. Accolade: Intermediate copying for reverse engineering was fair use if the final product was non-infringing.

These initial "results" demonstrate a judicial inclination towards finding AI training data uses transformative, but also highlight significant differences in how judges weigh factors like the source of data and the potential for indirect market harm. The lack of empirical evidence for market dilution proved critical in Kadrey, but the theory itself remains a potent signal for future litigation.

Practical Implications

▶ Watch: Authors Guild v. Google Books fair use precedent. (6:00)

The ongoing legal battles surrounding generative AI and copyright law have profound practical implications for a wide array of stakeholders, from AI developers to individual creators and even non-profit researchers.

For AI developers and infrastructure teams, the current landscape is fraught with significant uncertainty. If courts ultimately rule against fair use, the financial remedies are "extraordinarily generous." This includes:

  • Lost licensing revenues: Despite judges in Bartz and Kadrey rejecting a right to license training data, other judges might rule differently.
  • Defendants' profits attributable to infringement: This could be a colossal sum, given the projected revenues of AI companies.
  • Statutory damages: A minimum of $750 per infringed work, potentially escalating to $150,000 per work. With billions of works ingested, this could lead to damages in the "trillions of dollars."
  • Injunctions: Courts could halt further use of infringing models.
  • Model destruction: Some complaints explicitly ask for the "impoundment, destruction, or otherwise disposal of infringing materials," raising the specter of models being decommissioned. While discretionary, this is a severe potential outcome.

This high-stakes environment means AI companies must navigate a complex risk assessment, potentially influencing data acquisition strategies, model development, and even geographic deployment (regulatory arbitrage).

For model builders and deployers, the divergence in judicial opinions (e.g., on pirated books and market dilution) means that what might be considered fair use in one jurisdiction or under one judge's interpretation might not be elsewhere. This complicates compliance and legal strategy, especially for global operations. The emphasis on output filters (as mentioned by Meta) suggests that technical safeguards to prevent verbatim reproduction or substantial similarity in outputs will become increasingly important as a legal defense.

For practitioners and creators, the situation is a mixed bag. On one hand, the lawsuits represent an effort to secure fair compensation and protect livelihoods against perceived exploitation. If successful, this could lead to new revenue streams for creators, potentially through collective licensing schemes or direct payments. On the other hand, the legal process is slow and expensive, and the outcomes are highly uncertain. The "market dilution" theory, if it gains traction, could fundamentally alter how market harm is assessed, expanding the scope of potential claims.

International approaches offer potential models and challenges:

  • Opt-out mechanisms (EU): Give copyright owners some leverage to negotiate licenses, but also create administrative burdens.
  • Collective licensing: Widely used in other countries, this could offer a solution where AI companies pay into a fund that then distributes royalties to creators. However, practical problems abound, including the global nature of the internet, the sheer volume of data, and the difficulty of attributing value. The EU is actively exploring mandatory collective licensing.

Finally, non-profit researchers face spillover effects. While their uses are generally more likely to be deemed fair (due to non-commerciality and scholarly purpose), the current high-profile cases focus on big tech. Rulings against commercial AI could inadvertently narrow fair use interpretations, impacting academic research. Samuelson urged the technical community, including researchers, to actively engage in policy debates to ensure their interests are represented, advocating for a balance that promotes progress for the public good without undermining creators. The lack of clear disclosure requirements for training data further complicates scrutiny and accountability.

Key Takeaways

  • Copyright is a Major Regulatory Force: Generative AI's reliance on vast datasets has thrust copyright law, particularly the fair use doctrine, into the forefront of AI regulation.
  • High Stakes and Uncertainty: Over 44 lawsuits in the US, with potential damages in the trillions, create significant legal uncertainty for AI developers, likely lasting for years.
  • Transformative Use as Key Defense: AI companies primarily argue their use of copyrighted works for training is "highly transformative" because it serves a fundamentally different, non-expressive purpose (data extraction for model building).
  • Judicial Divergence on Key Factors: US judges have shown differing views on the significance of using pirated works for training and the novel "market dilution" theory (whether AI outputs can indirectly harm creators' markets).
  • Generous Remedies for Infringement: If found liable, AI companies face potentially massive financial penalties, injunctions, and even the destruction of their models.
  • Global Patchwork of Laws: International approaches vary, with some countries (e.g., Japan, Israel) enacting broad exceptions for AI, while others (e.g., EU) explore opt-out mechanisms or collective licensing, creating regulatory arbitrage opportunities.
  • Call for Engagement: The technical community is urged to actively participate in policy debates to shape the future of AI regulation and ensure a balance between innovation and creator rights.

About the Speaker(s)

Pamela Samuelson is a highly respected legal scholar, specializing in intellectual property law, particularly copyright and cyberlaw. Her detailed and nuanced analysis of the complex interplay between emerging technologies like generative AI and established legal frameworks demonstrates deep expertise in the field. As evidenced by her passionate call for the technical community to engage in policy debates, she is not only an academic but also an advocate for informed public discourse on these critical issues. Her ability to translate intricate legal concepts into understandable terms for a technical audience underscores her role as a vital bridge between the legal and technological worlds.

Reviews

Maya Iyer (Theoretical ML Researcher) — SOLID

Pamela Samuelson delivers a technically competent, well-structured survey of the current US copyright litigation landscape as it bears on generative AI training data. The talk is clear, practically relevant, and unusually valuable for an ML audience that is often underexposed to legal first principles. It is, however, fundamentally a legal survey talk rather than a research contribution — it synthesizes existing doctrine and ongoing cases but does not advance a novel legal thesis or produce a result the field did not have before. For ICML specifically, this earns its place as an invited talk: the audience needs this information. But evaluated as a contribution rather than a briefing, the…

Chen Zhao (Applied ML Researcher & Empiricist) — SOLID

Pamela Samuelson delivers a competent and well-structured orientation on copyright law and generative AI training data, targeted at an ML audience that largely lacks legal grounding. The talk surfaces genuinely important uncertainty — damages in the trillions, divergent rulings on pirated training data, the embryonic 'market dilution' theory — and does so with appropriate precision about what is settled versus contested. For a keynote-style invited talk at ICML, this is useful infrastructure knowledge. It is not, however, an empirical contribution in the ML sense, and the framework I use for rating research rigor maps awkwardly onto legal scholarship. Rated as a conference talk for ML…

→ Top-rated talks at International Conference on Machine Learning 2025

All talks from International Conference on Machine Learning 2025