Judge William Alsup ruled that Anthropic’s use of copyrighted texts to train its LLMs does not constitute infringement, even though the company was ordered to pay $1.5 billion to a group of authors. The penalty stemmed from Anthropic’s acquisition of the works through illegal shadow‑library sites, not from the training process itself. The judge likened an LLM’s ingestion of trillions of tokens to a writer studying literature, arguing the model creates new expression rather than copying existing works.
Intellectual‑property attorney Cathy Gellis said the decision treats model training as akin to reading a copyrighted work, which under current law does not trigger the copying element required for infringement. She noted the $1.5 billion fine is small relative to Anthropic’s projected $200 billion annual revenue by 2028. Another lawyer, Jason Henderson, warned that copyright statutes dating from 1976 leave courts to apply outdated principles to AI training, creating legal uncertainty for creators and developers.
- Training LLMs on copyrighted text is non‑infringing when obtained legally.
- $1.5 billion fine resulted from using pirated shadow‑library sources, not from the training act.
- Judge compared token ingestion to a writer’s study of literature, stressing transformation over replication.
- Experts say the ruling reflects a “reading” view of copyright, highlighting the 1976 statute’s mismatch with AI‑scale data use.
Why this matters
The decision suggests courts may treat large‑scale token ingestion as permissible “reading” under current copyright law, but the fine for using pirated sources shows that acquisition legality remains a concrete liability risk for AI developers.
Share this article
Found this insightful? Share it with your community on Reddit, X, or copy the link.
