Governance & ComplianceAI Products & Platforms

Same Book, Fair Use for Training — Why a Pirated Pipeline Costs $1.5 Billion

In Bartz v. Anthropic, multiple authors and rights holders alleged that the AI company Anthropic downloaded millions of books from pirate sources such as LibGen and PiLiMi and built an internal research library on its servers for AI development. On July 20, 2026, the court issued its Final Approval Order and Judgment (Final Approval Order and Judgment, Dkt. 680 (2026-07-20)), establishing a $1.5 billion non-reversionary settlement fund to resolve the class action. Of this massive fund, the court approved over $100 million in plaintiffs’ attorneys’ fees.

The settlement outcome surprised the industry. Prior to it, the court record had leaned toward finding that reading books into a large model for training qualified as fair use under U.S. copyright law, because such computation aims to learn linguistic patterns and does not undermine the commercial market for the original works. If the core training computation could be fair use, why did Anthropic agree to pay $1.5 billion to settle?

To understand this apparent contradiction, one must trace the flow of data through the technical pipeline. From downloading books, to retaining them on servers, to loading them into GPU memory for training, the same digital copy triggers different acts of reproduction at different stages. The court did not rule the entire LLM development process legal or illegal, nor did it judicially separate a compliant path from an infringing one. It simply assessed legal liability for different stages of the pipeline separately: first, the acquisition and retention of source files; second, loading those files into GPU memory for training.

What Happened in This Case, and Why Was the Outcome Surprising?

We can trace the lifecycle of a book through the R&D system. From the system downloading a file to a server, to finally loading it into GPU memory to update parameters, the same digital copy undergoes multiple acts of reproduction from a technical standpoint.

On June 23, 2025, the court issued its ruling (Order on Fair Use, Dkt. 231 (2025-06-23)) evaluating the training behavior of loading books into GPU memory to update model parameters. The judge was inclined to favor Anthropic, reasoning that machine training aims to extract statistical patterns — it does not consume a novel’s storyline the way a human reader does — and therefore qualifies as fair use.

But before training ever took place, the digital copy of the novel had long been sitting on Anthropic’s servers. Regarding the act of internally retaining pirated books for an extended period, the court declined to grant immunity before trial and denied Anthropic’s motion for summary judgment on fair use. This meant that, had the litigation proceeded, a jury would have decided whether internally retaining these books constituted infringement and, if so, what damages to award. Before a jury could reach a verdict, Anthropic paid a fixed fund to extinguish this clear, historic class-action risk. The essence of this deal: use a fixed sum to buy out the unresolved historical litigation risk that hung before a jury verdict. This was a negotiated risk settlement — it was neither an admission by Anthropic that it lost on this issue, nor a forward-looking general license, nor a declaration by the court or the agreement that the model itself was clean. To see this distinction clearly, we need to follow the full technical pipeline that data passes through inside the system.

Technically, this is one continuous pipeline. Let us follow the journey of a single book through the R&D system:

A single book copy from acquisition to retention to training: legal outcomes for different copying purposes in Bartz

First, the system downloads the digital file of the novel from publicly accessible sources such as LibGen or PiLiMi.

Second, the system stores the downloaded file locally, pooling it with other books into a long-term internal book repository. Court filings refer to this repository as the central library.

Third, when model training begins, the system retrieves the book from the central library, loads it into GPU memory for computation, and updates parameters.

Across the entire pipeline, the same file faces different legal assessments for the acts of reproduction occurring at different stages.

First, the computational copy loaded into GPU memory. The act of loading data into GPU memory to extract statistical patterns forms the cornerstone of the fair use determination.

But the central library sitting statically on disk in the preceding stage is an entirely different matter. Retaining pirated books in their original format on servers long-term, without any transformative new purpose, amounts to simple reproduction and storage. Regarding this act of retention, the court noted in its June 2025 ruling that it could not grant a fair use defense outright before trial.

The court’s refusal to grant immunity before trial does not mean it found the repository infringing. Since the case subsequently settled and the scheduled jury trial was cancelled, the court never rendered a final legal determination on the internal repository. But for AI teams, the intermediate retention step on disk did become the single greatest compliance vulnerability in the entire pipeline.

What Exactly Did $1.5 Billion Buy — and What Did It Not Buy?

Under the class action system, the compliance vulnerability at the intermediate stage eventually ballooned into enormous financial risk. The $1.5 billion settlement amount targets precisely the historical litigation risk accumulated at this stage.

To understand this figure, one must be familiar with the statutory damages regime unique to U.S. copyright law. Under this regime, plaintiffs need not prove actual economic losses; the court determines damages directly based on the number of infringed works. Under the U.S. Copyright Act (17 U.S.C. § 504(c)), the standard statutory damages range is $750 to $30,000 per work; if willful infringement is found, the maximum rises to $150,000.

In Bartz v. Anthropic, the plaintiffs obtained class certification on July 17, 2025 (Order on Class Certification, Dkt. 244 (2025-07-17)), allowing hundreds of thousands of rights holders to collectively pursue claims against Anthropic. Although the court denied certification for the Books3 and scanned-book classes, it certified a class covering books obtained from LibGen and PiLiMi.

The parties ultimately compiled a list of copyrighted works eligible for settlement — the Works List — totaling 482,460 works. Had the litigation continued, a jury would have decided whether retaining these books in the central library constituted infringement and determined the damages amount. If infringement were found, even at the statutory floor of $750 per work, total damages would exceed $360 million; if willful infringement were found, the ceiling could surpass $72 billion. To extinguish this clear, historic class-action risk before a jury could reach a verdict, Anthropic paid $1.5 billion to establish a fixed fund. The essence of the transaction: use a fixed payment to buy out the enormous, unresolved litigation risk that hung before a jury verdict. This was a negotiated risk settlement between the parties — it does not mean Anthropic admitted to losing on the retention issue, nor is it a forward-looking general license, nor does it represent the court or the agreement declaring the model itself fully clean.

Of the $1.5 billion non-reversionary fund, the court approved $101,561,111 in plaintiffs’ attorneys’ fees, with the remaining net proceeds distributed to rights holders. As of April 16, 2026, 91.3% of the works on the list had submitted claims. Dividing $1.5 billion across 482,460 works yields a rough nominal figure of approximately $3,109 per work. But it must be emphasized that this is a gross figure including attorneys’ fees and administrative costs; it is neither a standard copyright licensing fee nor can it serve as a reference price for future training authorization. Although the court issued its Final Approval Order and Judgment on July 20, 2026, the fund distribution, release of liability, and document destruction remain subject to the Effective Date as defined in the settlement agreement and will take effect and proceed gradually.

So, does paying this massive settlement sum mean that the resulting trained model is now fully compliant?

The answer is no. Neither the court’s rulings nor the settlement agreement itself issued a compliance certificate for the model.

First, the court’s inclination that model training may be fair use does not amount to a declaration that the resulting model weights are non-infringing, nor does it retroactively legalize the original source file repository. Likewise, the settlement agreement makes no determination about whether the model weights constitute infringement.

Second, while the settlement agreement (Class Action Settlement Agreement) includes a representation by Anthropic that its publicly released commercial models’ training sets did not contain LibGen or PiLiMi data, this is merely a unilateral representation and warranty by the contracting party — not a judicial fact investigated and found by the court.

Third, the settlement is, in substance, a payment to eliminate a specific historical litigation risk — not an all-purpose pardon. It releases only claims of reproduction infringement concerning the specific works on the Works List, during the specific time period, at the model training stage (the input side). Claims for output-side infringement (e.g., a user prompting the model to generate passages substantially similar to an original book), as well as future data use after the settlement’s effective date, are all explicitly excluded from the scope of the release.

Furthermore, the settlement agreement requires Anthropic to destroy the original books downloaded from LibGen and PiLiMi, as well as all local copies of those downloads (except copies retained as evidence to satisfy litigation preservation obligations), but does not require deletion or destruction of the already-trained model weights.

Settlement boundary of the $1.5 billion: the Works List, scope of release, and explicit exclusions

What Should AI Data Systems Be Able to Prove Going Forward?

The outcome of this litigation forces the AI industry to reexamine data compliance. The safety boundary is no longer an abstract courtroom debate — it has become a company’s ability to prove its compliance at the engineering level.

In the post-settlement era, LLM teams can no longer audit their data at the crude level of simply recording dataset names. A competent data management system must be able to prove the following at the engineering level:

First, acquisition channel. The system must be able to precisely trace the specific source of every file, ensuring the team can at any time identify and quarantine files that carry channel risk.

Second, purpose of copying. The system must clearly distinguish temporary in-memory compute caches from long-term static storage copies, preventing compliance risks arising from conflation of purposes.

Third, retention status. For temporary files that need not be retained long-term, the system should establish automated cleanup mechanisms to prevent data from lingering on servers.

Fourth, deletion and preservation status. If a rights holder issues a takedown request, or a court imposes a litigation preservation order, the system must be able to precisely locate, isolate, and destroy specific data while generating tamper-resistant audit logs.

For LLM development teams, upgrading their data management systems is not merely a compliance suggestion but an engineering baseline that must be implemented. A future data pipeline cannot merely provide the name of a dataset — it must produce a complete chain of evidence, linking together acquisition channel, purpose of copying, retention status, and deletion status.