In Bartz v. Anthropic, multiple authors and rights
holders alleged that the AI company Anthropic downloaded millions of
books from pirate sources such as LibGen and
PiLiMi and built an internal research library on its
servers for AI development. On July 20, 2026, the court issued its Final
Approval Order and Judgment (Final
Approval Order and Judgment, Dkt. 680 (2026-07-20)), establishing a
$1.5 billion non-reversionary settlement fund to resolve the class
action. Of this massive fund, the court approved over $100 million in
plaintiffs’ attorneys’ fees.
The settlement outcome surprised the industry. Prior to it, the court record had leaned toward finding that reading books into a large model for training qualified as fair use under U.S. copyright law, because such computation aims to learn linguistic patterns and does not undermine the commercial market for the original works. If the core training computation could be fair use, why did Anthropic agree to pay $1.5 billion to settle?
To understand this apparent contradiction, one must trace the flow of data through the technical pipeline. From downloading books, to retaining them on servers, to loading them into GPU memory for training, the same digital copy triggers different acts of reproduction at different stages. The court did not rule the entire LLM development process legal or illegal, nor did it judicially separate a compliant path from an infringing one. It simply assessed legal liability for different stages of the pipeline separately: first, the acquisition and retention of source files; second, loading those files into GPU memory for training.
We can trace the lifecycle of a book through the R&D system. From the system downloading a file to a server, to finally loading it into GPU memory to update parameters, the same digital copy undergoes multiple acts of reproduction from a technical standpoint.
On June 23, 2025, the court issued its ruling (Order on Fair Use, Dkt. 231 (2025-06-23)) evaluating the training behavior of loading books into GPU memory to update model parameters. The judge was inclined to favor Anthropic, reasoning that machine training aims to extract statistical patterns — it does not consume a novel’s storyline the way a human reader does — and therefore qualifies as fair use.
But before training ever took place, the digital copy of the novel had long been sitting on Anthropic’s servers. Regarding the act of internally retaining pirated books for an extended period, the court declined to grant immunity before trial and denied Anthropic’s motion for summary judgment on fair use. This meant that, had the litigation proceeded, a jury would have decided whether internally retaining these books constituted infringement and, if so, what damages to award. Before a jury could reach a verdict, Anthropic paid a fixed fund to extinguish this clear, historic class-action risk. The essence of this deal: use a fixed sum to buy out the unresolved historical litigation risk that hung before a jury verdict. This was a negotiated risk settlement — it was neither an admission by Anthropic that it lost on this issue, nor a forward-looking general license, nor a declaration by the court or the agreement that the model itself was clean. To see this distinction clearly, we need to follow the full technical pipeline that data passes through inside the system.
Technically, this is one continuous pipeline. Let us follow the journey of a single book through the R&D system:
First, the system downloads the digital file of the novel from
publicly accessible sources such as LibGen or
PiLiMi.
Second, the system stores the downloaded file locally, pooling it
with other books into a long-term internal book repository. Court
filings refer to this repository as the
central library.
Third, when model training begins, the system retrieves the book from
the central library, loads it into GPU memory for
computation, and updates parameters.
Across the entire pipeline, the same file faces different legal assessments for the acts of reproduction occurring at different stages.
First, the computational copy loaded into GPU memory. The act of loading data into GPU memory to extract statistical patterns forms the cornerstone of the fair use determination.
But the central library sitting statically on disk in
the preceding stage is an entirely different matter. Retaining pirated
books in their original format on servers long-term, without any
transformative new purpose, amounts to simple reproduction and storage.
Regarding this act of retention, the court noted in its June 2025 ruling
that it could not grant a fair use defense outright before trial.
The court’s refusal to grant immunity before trial does not mean it found the repository infringing. Since the case subsequently settled and the scheduled jury trial was cancelled, the court never rendered a final legal determination on the internal repository. But for AI teams, the intermediate retention step on disk did become the single greatest compliance vulnerability in the entire pipeline.
Under the class action system, the compliance vulnerability at the intermediate stage eventually ballooned into enormous financial risk. The $1.5 billion settlement amount targets precisely the historical litigation risk accumulated at this stage.
To understand this figure, one must be familiar with the statutory damages regime unique to U.S. copyright law. Under this regime, plaintiffs need not prove actual economic losses; the court determines damages directly based on the number of infringed works. Under the U.S. Copyright Act (17 U.S.C. § 504(c)), the standard statutory damages range is $750 to $30,000 per work; if willful infringement is found, the maximum rises to $150,000.
In Bartz v. Anthropic, the plaintiffs obtained class
certification on July 17, 2025 (Order
on Class Certification, Dkt. 244 (2025-07-17)), allowing hundreds of
thousands of rights holders to collectively pursue claims against
Anthropic. Although the court denied certification for the Books3 and
scanned-book classes, it certified a class covering books obtained from
LibGen and PiLiMi.
The parties ultimately compiled a list of copyrighted works eligible
for settlement — the Works List — totaling 482,460 works.
Had the litigation continued, a jury would have decided whether
retaining these books in the central library constituted
infringement and determined the damages amount. If infringement were
found, even at the statutory floor of $750 per work, total damages would
exceed $360 million; if willful infringement were found, the ceiling
could surpass $72 billion. To extinguish this clear, historic
class-action risk before a jury could reach a verdict, Anthropic paid
$1.5 billion to establish a fixed fund. The essence of the transaction:
use a fixed payment to buy out the enormous, unresolved litigation risk
that hung before a jury verdict. This was a negotiated risk settlement
between the parties — it does not mean Anthropic admitted to losing on
the retention issue, nor is it a forward-looking general license, nor
does it represent the court or the agreement declaring the model itself
fully clean.
Of the $1.5 billion non-reversionary fund, the court approved $101,561,111 in plaintiffs’ attorneys’ fees, with the remaining net proceeds distributed to rights holders. As of April 16, 2026, 91.3% of the works on the list had submitted claims. Dividing $1.5 billion across 482,460 works yields a rough nominal figure of approximately $3,109 per work. But it must be emphasized that this is a gross figure including attorneys’ fees and administrative costs; it is neither a standard copyright licensing fee nor can it serve as a reference price for future training authorization. Although the court issued its Final Approval Order and Judgment on July 20, 2026, the fund distribution, release of liability, and document destruction remain subject to the Effective Date as defined in the settlement agreement and will take effect and proceed gradually.
So, does paying this massive settlement sum mean that the resulting trained model is now fully compliant?
The answer is no. Neither the court’s rulings nor the settlement agreement itself issued a compliance certificate for the model.
First, the court’s inclination that model training may be fair use does not amount to a declaration that the resulting model weights are non-infringing, nor does it retroactively legalize the original source file repository. Likewise, the settlement agreement makes no determination about whether the model weights constitute infringement.
Second, while the settlement agreement (Class
Action Settlement Agreement) includes a representation by Anthropic
that its publicly released commercial models’ training sets did not
contain LibGen or PiLiMi data, this is merely
a unilateral representation and warranty by the contracting party — not
a judicial fact investigated and found by the court.
Third, the settlement is, in substance, a payment to eliminate a
specific historical litigation risk — not an all-purpose pardon. It
releases only claims of reproduction infringement concerning the
specific works on the Works List, during the specific time
period, at the model training stage (the input side). Claims for
output-side infringement (e.g., a user prompting the model to generate
passages substantially similar to an original book), as well as future
data use after the settlement’s effective date, are all explicitly
excluded from the scope of the release.
Furthermore, the settlement agreement requires Anthropic to destroy
the original books downloaded from LibGen and
PiLiMi, as well as all local copies of those downloads
(except copies retained as evidence to satisfy litigation preservation
obligations), but does not require deletion or destruction of the
already-trained model weights.
The outcome of this litigation forces the AI industry to reexamine data compliance. The safety boundary is no longer an abstract courtroom debate — it has become a company’s ability to prove its compliance at the engineering level.
In the post-settlement era, LLM teams can no longer audit their data at the crude level of simply recording dataset names. A competent data management system must be able to prove the following at the engineering level:
First, acquisition channel. The system must be able to precisely trace the specific source of every file, ensuring the team can at any time identify and quarantine files that carry channel risk.
Second, purpose of copying. The system must clearly distinguish temporary in-memory compute caches from long-term static storage copies, preventing compliance risks arising from conflation of purposes.
Third, retention status. For temporary files that need not be retained long-term, the system should establish automated cleanup mechanisms to prevent data from lingering on servers.
Fourth, deletion and preservation status. If a rights holder issues a takedown request, or a court imposes a litigation preservation order, the system must be able to precisely locate, isolate, and destroy specific data while generating tamper-resistant audit logs.
For LLM development teams, upgrading their data management systems is not merely a compliance suggestion but an engineering baseline that must be implemented. A future data pipeline cannot merely provide the name of a dataset — it must produce a complete chain of evidence, linking together acquisition channel, purpose of copying, retention status, and deletion status.