Back to Newsroom

Is it legal to train AI models on copyrighted books? It’s complicated

By Modelverse Editorial·August 23, 2026·2 min read
Is it legal to train AI models on copyrighted books? It’s complicated

Judge William Alsup ruled that Anthropic’s use of copyrighted texts to train its LLMs does not constitute infringement, even though the company was ordered to pay $1.5 billion to a group of authors. The penalty stemmed from Anthropic’s acquisition of the works through illegal shadow‑library sites, not from the training process itself. The judge likened an LLM’s ingestion of trillions of tokens to a writer studying literature, arguing the model creates new expression rather than copying existing works.

Intellectual‑property attorney Cathy Gellis said the decision treats model training as akin to reading a copyrighted work, which under current law does not trigger the copying element required for infringement. She noted the $1.5 billion fine is small relative to Anthropic’s projected $200 billion annual revenue by 2028. Another lawyer, Jason Henderson, warned that copyright statutes dating from 1976 leave courts to apply outdated principles to AI training, creating legal uncertainty for creators and developers.

  • Training LLMs on copyrighted text is non‑infringing when obtained legally.
  • $1.5 billion fine resulted from using pirated shadow‑library sources, not from the training act.
  • Judge compared token ingestion to a writer’s study of literature, stressing transformation over replication.
  • Experts say the ruling reflects a “reading” view of copyright, highlighting the 1976 statute’s mismatch with AI‑scale data use.

Why this matters

The decision suggests courts may treat large‑scale token ingestion as permissible “reading” under current copyright law, but the fine for using pirated sources shows that acquisition legality remains a concrete liability risk for AI developers.

Share this article

Found this insightful? Share it with your community on Reddit, X, or copy the link.

ai-newsbrieftechcrunch-ai

Footnotes & Primary References

Related content

Building an End-to-End Document Intelligence Pipeline with deepDoctection

Build an end-to-end document intelligence pipeline with deepDoctection. This tutorial covers configuring layout analysis, DocTR OCR, and table extraction, while demonstrating how t...

Read article

Vercel Introduces 'Is Agentic', a Free Agent-Readiness Scoring Tool That Audits Public Websites Using Ora's 100+ Checks

Vercel and Ora launched Is Agentic, a free audit scoring website readiness for AI agents across 118 checks. The post Vercel Introduces 'Is Agentic', a Free Agent-Readiness Scoring ...

Read article

Ox Alpha Lands on OpenRouter as a Free 1M-Context Stealth Model for Coding and Agentic Work

Ox Alpha is a new anonymous reasoning model on OpenRouter with a 1M-token context window, multimodal inputs, tool calling and free preview pricing — aimed at coding and long-running AI agents.

Read article