T1#regulation#ethics#market
Bartz v. Anthropic — Training an LLM Is Fair Use; Stocking a Library with Pirated Books Is Not

Metadata
- Date
- Decade
- 2020s
- Tier
- T1
- Sources
- 10
- Connections
- 02
- Tags
- #regulation#ethics#market
On 23 June 2025, Judge William Alsup of the United States District Court for the Northern District of California filed a 32-page "Order on Fair Use". Dozens of copyright suits against AI developers were already pending in US courts; this was the first to reach a substantive ruling on whether training a large language model on copyrighted books is a fair use.
The plaintiffs were three authors — Andrea Bartz, Charles Graeber and Kirk Wallace Johnson — together with two corporate entities two of them had set up to market their work. They filed on 19 August 2024. The defendant, Anthropic PBC, had used their books to train the models behind Claude.
One case, three answers
Alsup declined to treat Anthropic's conduct as a single use. He assessed fair use separately for each purpose the copies served, and that decomposition is the whole architecture of the opinion.
| Conduct | Holding |
|---|---|
| Using the books to train LLMs | Fair use |
| Digitising lawfully purchased print copies | Fair use, on different reasoning |
| Downloading pirated copies to build a central library | Not fair use — set for trial |
On training, the order states that the use of the books at issue to train Claude and its precursors "was exceedingly transformative and was a fair use under Section 107 of the Copyright Act". Elsewhere it puts the point more bluntly: the purpose and character of using works to train LLMs was "transformative — spectacularly so".
The digitisation holding rests on something else entirely. Anthropic bought millions of print books, often used, had a service provider tear off the bindings, cut the pages to size and scan them, and discarded the paper originals. That, the court reasoned, was not transformation but replacement: a more convenient, space-saving, searchable copy substituted for one the company already owned, "without adding new copies, creating new works, or redistributing existing copies."
Then the third bucket. In January or February 2021, co-founder Ben Mann downloaded Books3, a corpus of 196,640 books he knew had been assembled from unauthorised copies. In June 2021 he pulled at least five million books from Library Genesis. In July 2022 Anthropic took at least two million more from the Pirate Library Mirror. The court found the company had pirated over seven million copies of books. Building a permanent, general-purpose library — "all the books in the world", retained "forever", in the phrases the order quotes from Anthropic's own documents — was not itself a fair use that excused how the books were obtained.
"Training schoolchildren to write well"
The market-harm analysis is where the opinion reaches furthest.
The authors argued that training LLMs will produce an explosion of works competing with theirs. The court assumed that was true and held it irrelevant: their complaint, Alsup wrote, "is no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works. This is not the kind of competitive or creative displacement that concerns the Copyright Act. The Act seeks to advance original works of authorship, not to protect authors against competition."
The authors also argued that Anthropic had displaced an emerging market for licensing books specifically for AI training. The court accepted that such a market could develop — and still held that a market for that use is not one the Copyright Act entitles authors to exploit.
What the order pointedly does not decide is worth as much as what it does. The authors never alleged that any Claude output infringed their works. Alsup says so twice, notes that the record shows the opposite, and adds that if it were otherwise "this would be a different case" and that the authors remain free to bring it should the facts develop.
From ruling to US$1.5 billion
The pirated copies went toward trial. On 17 July 2025 Alsup certified a LibGen & PiLiMi Pirated Books Class — and denied certification of both a Books3 class and a scanned-books class. That detail is routinely lost in summaries of the case: the certified class was defined by two specific pirate corpora, not by everything Anthropic ever downloaded.
US copyright law allows statutory damages up to $150,000 per work for wilful infringement. Multiplied against a finding of seven million pirated copies, the theoretical exposure set the terms of everything that followed.
The parties signed a term sheet on 25 August 2025 and a settlement agreement on 5 September: US$1.5 billion, non-reversionary, paid in four instalments — $300 million within five business days of preliminary approval, $300 million within five business days of final approval, $450 million within twelve months of preliminary approval and $450 million within twenty-four, each of the last two accruing interest at the three-month Treasury bill rate. The agreement also obliges Anthropic to destroy the original files torrented from LibGen and PiLiMi, and copies originating from them, within thirty days of final judgment, and to certify the deletion in writing to class counsel. Scanned copies are explicitly excluded from that obligation.
After hearings on 8 and 25 September, Alsup granted preliminary approval at the second. His written opinion of 17 October 2025 called the deal "the largest copyright class action settlement in history".
Final approval came on 20 July 2026 from Judge Araceli Martínez-Olguín, the case having been reassigned after Alsup's retirement at the end of 2025. The Works List holds 482,460 works; as of 16 April 2026, 91.3 per cent had been claimed. The per-work payment is about $3,000 — four times the $750 statutory minimum for ordinary infringement. Class counsel asked for 12.5 per cent of the fund ($187.5 million) and were awarded $101,561,111, roughly 6.8 per cent. Service awards to the three class representatives were cut from the requested $50,000 each to $15,000.
Paragraph 1.29 and its three limits
The scope of the release is the most misreported part of the settlement, and it is stated precisely in the agreement.
Paragraph 1.29 releases claims arising from Anthropic's past torrenting, scanning, retention and use — including training — of works on the Works List. It then adds three limits in the same breath: "no claims based on the output of AI models are released"; there is no release of future reproduction, distribution or derivative works; and the release "does not constitute, in any respect, a license to torrent, scan, or train AI models on any copyrighted works". Finally: "Released Claims do not extend to any activity or conduct that occurs or occurred after August 25, 2025." That date is not arbitrary — it is the day the parties signed the term sheet.
The final approval order drew the same line when overruling objections, holding that class members give up claims about past inputs but "do not give up claims about past AI outputs, nor claims of any kind about future conduct (on or after August 25, 2025)."
Anthropic's deputy general counsel Aparna Sridhar told NPR: "Today's settlement, if approved, will resolve the plaintiffs' remaining legacy claims." Legacy — that is, a question about provenance, not about the model.
Not precedent, but a price
The reach of Bartz is narrower than the headlines. It is a district court ruling, not an appellate one, and because the case settled, the piracy question never produced a judgment that binds anyone. But the line Alsup drew was legible: the training is not the problem; the sourcing is.
Two days later, on 25 June 2025, Judge Vince Chhabria of the same district granted summary judgment to Meta in Kadrey v. Meta — also fair use, but reached differently, on the ground that the plaintiffs had failed to build a record of market dilution. Four months earlier, Judge Stephanos Bibas in Delaware had rejected fair use in Thomson Reuters v. ROSS Intelligence, a case about a legal-research tool rather than a generative model. Three rulings in 2025, three routes to an answer.
For an industry that had spent three years treating training corpora as an engineering problem, the $3,000-per-work figure was the first time the question carried a price. It is not precedent. It became a starting point for negotiation anyway.
Sources
Last updated: