Anthropic Destroyed Millions of Books to Build Claude’s Training Library

The physical destruction was part industrial efficiency, part legal strategy and part a wider battle over ownership of knowledge.

SAN FRANCISCO, UNITED STATES

Anthropic secretly purchased, dismantled and scanned millions of printed books to create a high-quality digital library for training its Claude artificial-intelligence models. The internal initiative, known as Project Panama, emerged through documents released during a copyright lawsuit brought by authors against the company.

The process was deliberately destructive. Contractors used hydraulic machinery to remove the bindings, separate the pages and feed them through industrial scanners. The remaining paper was sent for recycling rather than returned to libraries or sold. Anthropic sought suppliers capable of processing between 500,000 and two million books within six months.

Books offered professionally edited human writing that Anthropic considered more valuable than large quantities of informal internet content. Older publications were particularly useful because they predated the rapid expansion of AI-generated material. Training repeatedly on synthetic text can reinforce errors and reduce linguistic diversity, a potential degradation commonly described as model collapse.

Destroying each physical copy also supported Anthropic’s legal position. The company argued that converting a lawfully purchased book into an internal digital format did not create an additional market copy because the original was eliminated. A federal judge agreed that this specific format-conversion process and the subsequent use of the text for model training qualified as fair use under the circumstances examined.

That ruling did not legalise every method Anthropic used to obtain books. Court records also showed that the company had previously downloaded millions of works from pirate repositories including LibGen and Pirate Library Mirror. Anthropic later agreed to a $1.5 billion settlement covering approximately 500,000 eligible titles connected to those unauthorised acquisitions.

The distinction is central to the case. The court treated training on legally acquired books as transformative, while leaving Anthropic exposed for obtaining other copies through piracy. Project Panama therefore reveals more than an unusual scanning operation. It shows how artificial-intelligence companies transformed books into strategic data while copyright law struggled to distinguish lawful analysis, unauthorised copying and the economic rights of authors.

Detrás de cada dato, la intención. / Behind every data point, the intention.

Related posts

Justin and Hailey Bieber Face Renewed Marriage Speculation

Emmy Giving Suite Turns Celebrity Gifts Into Strategic Marketing

Miley Cyrus Becomes “Miley” for a New Musical Era