All the Books in the World
Libraries worth of physical books should not be fuel for industrial-level destructive scanning by AI companies.
Anthropic reportedly spent tens of millions of dollars buying millions of physical books, slicing off their bindings with machines, separating their pages for high-speed scanning, and sending the remains for recycling. The resulting files went into a private research database used to support development of Claude and related models. The operation had an internal name, “Project Panama,” and an internal description that was even more revealing: an effort to “destructively scan all the books in the world.” (washingtonpost.com)
The precise total has not been made public. Court materials establish that Anthropic purchased millions of print books, frequently used copies, and that its vendors stripped bindings, cut pages, scanned them into searchable PDFs, and discarded the originals. A vendor proposal described capacity to convert between 500,000 and 2 million books over six months. (fbm.com)
- Anthropic did not merely collect text from the internet. Its reported operation converted the book trade into a physical input pipeline for a frontier model developer: bulk acquisition, warehouse processing, cutting, scanning, OCR, metadata, private storage, and model training. The clean interface of a chatbot rests on a decidedly material system.
- The company’s purchase-and-scan program must be distinguished from the separate piracy dispute. In the June 23, 2025 ruling, the court found that Anthropic’s conversion of lawfully purchased print books into internal digital copies was fair use on the record before it. The court treated that program separately from Anthropic’s acquisition and retention of more than seven million pirated book copies from shadow libraries. The later $1.5 billion settlement addressed claims involving those pirated acquisitions, not the purchased physical books.
- Legal permission is not a preservation ethic. The court record does not indicate that Anthropic targeted rare books. But the cultural problem does not depend on a claim that every volume was rare. Millions of common physical books are still durable, independently ownable artifacts. They can remain on shelves, circulate through families, be loaned privately, resold, annotated, carried across borders, and read without a login, subscription, platform policy, or provider’s permission.
- This is part of a broader data-supply problem. As frontier labs seek large bodies of human-authored text that predate the spread of synthetic content, books have become a procurement target. The current book market is responding with brokers and bulk sourcing services that can locate, aggregate, and move physical volumes at scale. The result is an emerging training-data supply chain whose incentive is extraction, speed, and exclusivity, not public preservation. (tomshardware.com)
Orthogonal Take
A book is not just training data awaiting extraction. It is a physical record of human thought. It is stable in a way a model output is not. It can be privately owned, passed from hand to hand, opened without an account, preserved outside corporate infrastructure, and read in the same form decades after it was published.
That matters because a model does not preserve a book as a book. It does not preserve its edition, its sequence, its design, its physical history, its marginalia, or a reader’s independent right to return to the original. It absorbs patterns and produces new output through an opaque system whose training, safety policies, ranking logic, source availability, and answers may change at the discretion of its operator.
There is value in digitization, search, accessibility, and tools that help people navigate knowledge. The objection is not that machines can scan and digitize books. The objection is that we are increasingly treating a corporation’s private, processed remainder as an adequate replacement for the durable cultural artifact from which it was derived.
A library of books preserves plurality, contradiction, memory, and access to the source. A large language model transforms text into a probabilistic output.
Those are not the same civic function.
A society that devours its physical books by the millions in order to generate more mutable, filtered, and centrally controlled text is not a durable culture.
Note: This article was drafted with the assistance of AI and is for informational purposes only. It does not constitute legal or professional advice of any kind.