Anthropic's $1.5B Settlement and Thomson Reuters Case Clarify AI Training Copyright Law, Say TechCrunch, Gellis, and Henderson
On August 23, 2026, TechCrunch published an explainer detailing the complex legal status of training AI models on copyrighted books, drawing on insights from IP attorneys Cathy Gellis and Jason Henderson. This analysis, which discussed key rulings like the Anthropic $1.5 billion copyright settlement and the Thomson Reuters v. Ross Intelligence fair-use defeat, underscores how these decisions are shaping the entire AI industry.
Evolving Legal Interpretations for AI
Current copyright law, last updated in 1976, predates the advent of modern AI technologies. This gap has compelled courts to interpret existing statutes in novel ways to address the challenges posed by AI model training. The core issue revolves around whether the use of copyrighted works for training AI constitutes fair use or copyright infringement.
Key Rulings Shaping AI Copyright
Several pivotal legal cases have begun to establish precedents in this evolving area:
- Anthropic $1.5 Billion Copyright Settlement: This settlement, detailed in a TechCrunch report on September 5, 2025, specifically penalized the unlawful acquisition of copyrighted material from 'shadow libraries' rather than the act of AI training itself. Judge William Alsup's ruling found the training of AI models to be lawful, but the piracy of books for that training was deemed illegal.
- Thomson Reuters v. Ross Intelligence: In this case, the court determined that training an AI product that directly competes with the source material's creator does not qualify as fair use. This ruling suggests a limitation on how AI models can be trained when the output directly infringes on the market of the original copyrighted work.
- Thaler v. Perlmutter: This ruling established that works generated entirely by AI are not eligible for copyright protection. This decision has implications for the ownership and protection of content created solely by artificial intelligence systems.
Distinguishing Lawful Training from Infringement
The distinction between lawful AI training and copyright infringement often hinges on the method of data acquisition and the nature of the AI's output. While the act of training an AI model on publicly available data may be permissible, obtaining copyrighted materials through illicit means, such as pirated 'shadow libraries,' is clearly unlawful. Furthermore, if an AI model's output directly competes with and substitutes for the original copyrighted work, it may face legal challenges under fair use doctrines.
Implications for the AI Industry
The ongoing legal developments have significant implications for AI developers and companies. The need for robust legal analysis in constructing and utilizing large-scale datasets is increasingly recognized, as highlighted in a 2019 analysis by Mittelstadt. Companies developing AI models, such as those behind Anthropic's Claude, ChatGPT, or Gemini, must navigate these complexities to ensure their training practices comply with existing and emerging copyright laws. The outcomes of these cases will ultimately define the boundaries within which AI innovation can occur, influencing everything from data sourcing strategies to product development and market competition.
Conclusion
The legal framework for AI training on copyrighted content is still under construction, with courts improvising rules based on a 1976 copyright law. Key rulings have clarified that while AI training itself may be lawful, the illicit acquisition of training data and the creation of directly competing products are not. As the AI industry continues to evolve, developers must remain vigilant regarding these legal precedents to ensure ethical and lawful development practices.
Sources
- 1 Introduction
- Is it legal to train AI models on copyrighted books? It’s complicated | TechCrunch
- Screw the money -- Anthropic's $1.5B copyright settlement sucks for writers | TechCrunch
- Zoom knots itself a legal tangle over use of customer data for training AI models | TechCrunch
- It sure looks like OpenAI trained Sora on game content — and legal experts say that could be a problem | TechCrunch
Recommended AI tools
ChatGPT
Conversational AI
AI research, productivity, and conversation—smarter thinking, deeper insights.
Perplexity
Search & Discovery
Clear answers from reliable sources, powered by AI.
Claude
Conversational AI
Your trusted AI collaborator for coding, research, productivity, and enterprise challenges
OpenClaw AI Agent
Productivity & Collaboration
The AI that actually does things.
Cursor
Code Assistance
The AI code editor that understands your entire codebase
DeepSeek
Conversational AI
Efficient open-weight AI models for advanced reasoning and research
Was this article helpful?
Found outdated info or have suggestions? Send us a note.