Extract from Eureka Street
- Home
- Vol 36 No 16
- To build Claude, Anthropic needed to destroy millions of books
- Binoy Kampmark
- 20 August 2026
Last month, an investigation by 404 Media reported on the purchasing habits of Artificial Intelligence (AI) companies regarding old books (immediately, the issue of age crops up), identifying ISBNdb as a seminal figure in the field. The company, according to its own description, ‘gathers data from various public sources like libraries and merchants to compile a vast collection of unique book data searchable by ISBN, title, author or publisher.’ At present, it boasts 111,978,817 searchable books and offers clients somewhere between 1,000 and 1 million books per engagement. Older print publications are advertised as the purer sort, uncontaminated by the presence of generative-AI. ‘Print books from the pre-LLM era are structurally guaranteed to be free of this contamination,’ states an article published by the company. It remains unclear whether ISBNdb’s client list includes many of the destructive scanning persuasion, though the company is not oblivious to the problem.
ISBNdb by no means has a monopoly in this field. The Canadian company Zoom Books specialises in acquiring nonfiction and academic titles from the 1970s in bulk. These are then stored in German warehouses prior to their transport to the United States and Canada. It is there that AI companies then go to work scanning and destroying the hard copies. The practice struck Thomas Koch, press spokesperson for the German Publishers and Booksellers Association, as abominable not only because of its destructive dimension to the physical book, but also because of the violation of copyright arising from feeding LLMs ‘without consent and without payment.’
Anthropic’s activities came to light in documents filed in the Northern District of California case Bartz v Anthropic PBC, involving the authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson, all of whom alleged copyright infringement by the company. Of keen interest was the role of ‘Project Panama’, led by Tom Turvey, a former employee of Google and a central figure behind the creation of Google Books. The project, as the court order details, was intended to create ‘a central library of “all the books in the world” to retain “forever.”’ The AI firm could then mine the quarry for “various sets and subsets of digitized books to train various large language models under development to power its AI services.”
The destructive terminus for the scanned books was revealed in an internal memorandum from April 13, 2024. Authored by Turvey, it is unequivocal: ‘Project Panama is our effort to destructively scan all of the books in the world.’ The memorandum encourages company employees to use discretion when discussing the project. ‘Why use a codename? We use “soft codename” because we don’t want to be known that we are working on this. This document is available to all Anthropic employees, but you should avoid talking about it in public areas and the fact that we are working on this should not be shared with anyone outside Anthropic.’
Instead of seeking individual agreements with publishers to license copies for training AI, a process scoffingly dismissed as “legal/practice/business slog”, the team led by Turvey ‘emailed major book distributors and retailers about bulk purchasing their print copies for the AI firm’s “research library”.’ Millions of dollars were expended on printed books, even those in used condition. Retained service providers then went about their work: stripping the books of their bindings, adjusting the pages to size and digitising the copies. ‘Each print copy book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books).’ Paper originals were discarded. Not content with seeking hard copies, Anthropic also sought copies aplenty from the pirate-vendor market, including five million copies from LibGen (five million copies), two million items from Pirate Library Mirror (2 million) and Books 3 (approximately 183,000.
This orgy of physical destruction did not particularly interest Judge William Alsup. Of primary concern was whether copyright violations had taken place. Regarding the created ‘research library’, Judge Alsup found no issue with Anthropic’s conversion of each book copy’s format from print to digital. ‘Anthropic purchased its print copies fair and square.’ Insofar as the purchase and scanning of hard-copy books was concerned, the company’s LLMs failed to reproduce ‘to the public a given work’s creative elements, nor even one author’s identifiable expressive style’. Claude, the primary model in question, ‘outputted grammar, composition, and style that the underlying LLM distilled from thousands of works.’ But copyright did not cover ‘“method[s] of operation, concept[s], [or] principle[s]” “illustrated” [ ] or embodied in [a] work.’ The use of copyrighted works to train LLMs to generate the new text was ‘quintessentially transformative.’
'It gives free rein to companies to pursue generative AI development, a process that will, in time, transform the author not into an enlightened producer of knowledge but a beholden consumer, ever at the mercy of Claude and others of its ilk.'
The authors who filed the action had some joy regarding the pirated
copies. ‘The person who copies the textbook from a pirate site has
infringed already, full stop.’ Anthropic’s claim that using copies for a
central library was a case of fair use was untenable. The court further
rejected the company’s contention ‘that the use of copies for a central
library can be excused as fair use merely because some will eventually
be used to train LLMs.’
The judgment is apocalyptic for those who
believe in the sensuality of knowledge gained through a physical book.
With characteristic bluntness, Mike Masnick of Techdirt identified
the essence of the case. Alsup had ‘essentially created a roadmap that
validates legitimate AI training while drawing clear lines around what
crosses into infringement.’ Unfortunately, it gives free rein to
companies to pursue generative AI development, a process that will, in
time, transform the author not into an enlightened producer of knowledge
but a beholden consumer, ever at the mercy of Claude and others of its
ilk.
No comments:
Post a Comment