The Quiet Destruction of Our Literary Heritage: AI’s Insatiable Hunger for Rare Books
In an era defined by technological advancement, a troubling practice has emerged that strikes at the heart of human civilization: the systematic destruction of rare and irreplaceable physical books to feed the voracious appetite of artificial intelligence systems.
What was once the domain of preservationists and scholars, protecting texts that survived wars, fires, floods, and centuries of human handling, is now being undone in service of training large language models.
This is not hyperbole.
Reports from investigative outlets and court documents reveal a pipeline in which AI companies bulk-purchase books, slice off their spines for high-speed scanning, and shred or recycle the originals. Meanwhile, recent legal rulings have largely allowed the process to continue.
How Books Become AI Training Data
The mechanics are coldly efficient.
Companies such as ISBNdb, which describes itself as maintaining the world’s largest book database, have reportedly pivoted toward supplying AI firms with massive book orders. These orders can range from several thousand volumes to as many as one million books at a time.
Many buyers specifically seek editions published before 2022 because they predate the explosion of AI-generated content now contaminating digital information.
The buyers often remain anonymous, protected by non-disclosure agreements offered as a standard feature of the transaction.
Once the books arrive:
- Their bindings are removed.
- Their pages are separated.
- The content is scanned at industrial speed.
- The remaining paper is pulped, recycled, or discarded.
What survives is not a preserved book, but digital data absorbed into proprietary AI systems and often inaccessible to the public.
These Are Not Always Disposable Books
It would be easier to defend this practice if it involved only surplus copies of widely available books.
But reports suggest otherwise.
Booksellers have described volumes with very few surviving copies entering the same destructive pipeline. Some of these works may have endured for centuries, only to be reduced to training material for commercial AI systems.
Imagine an 18th-century botanical treatise containing hand-colored plates, unusual printing techniques, and handwritten annotations. Even if its words are scanned, much of its historical value cannot be captured through ordinary digitization.
Once the physical copy is shredded, the loss is permanent.
A popular novel can be reprinted. A website can be restored from a backup. A rare edition containing unique materials, provenance, or marginalia cannot be recreated after destruction.
Anthropic and “Project Panama”
At the center of this controversy stands Anthropic, the company behind the Claude family of AI models.
Court filings from the Bartz v. Anthropic case describe an internal initiative known as “Project Panama.” Its ambition was reportedly to acquire an enormous physical library, described internally as an effort to obtain “all the books in the world.”
Anthropic hired Tom Turvey, formerly head of partnerships for Google Books, to help lead the project.
Millions of used books were reportedly purchased. Their spines were removed, their pages were scanned, and the paper was recycled.
The objective was clear: build a vast, clean dataset of professionally written human knowledge without relying exclusively on an internet increasingly polluted by low-quality and AI-generated material.
The Court Ruling and the Fair Use Argument
In 2025, a federal judge in San Francisco ruled that using legally purchased books for AI training could qualify as fair use.
The reasoning was that AI training is transformative. The system does not merely republish the books. Instead, it uses them to develop new capabilities, such as answering questions, generating language, summarizing information, and assisting with research or coding.
The destruction of the original physical book also played a role in the legal reasoning. By removing the physical copy after scanning, the process was framed as limiting duplication rather than creating multiple competing copies.
From a copyright perspective, the ruling focused primarily on whether authors’ economic rights were being violated.
But that framework leaves a much larger question unanswered.
What happens to the cultural value of the physical book itself?
A Book Is More Than Its Words
Copyright law protects creative ownership, reproduction rights, and economic interests.
It does not adequately protect the historical significance of a physical object.
Rare books are not merely containers of text. They may also preserve:
- Handwritten notes and marginalia
- Ownership marks and signatures
- Unique bindings
- Paper quality and manufacturing techniques
- Printing errors and edition variations
- Evidence of censorship or revision
- Regional publishing history
- Clues about how earlier readers interpreted the work
A digital scan may capture the printed words, but it cannot always preserve the full material history of the object.
Destroying the book erases those layers permanently.
Why AI Companies Want Older Books
The incentives behind this practice are powerful.
Modern AI systems require enormous quantities of high-quality training data. Yet the internet is becoming increasingly contaminated by synthetic content generated by earlier AI systems.
This creates a dangerous feedback loop.
Models begin training on material produced by other models. Errors are repeated. Writing becomes more generic. Facts become distorted. Diversity of expression declines.
Physical books, especially those published before the large-scale adoption of generative AI, represent a finite source of edited, curated, and human-produced information.
They contain professional scholarship, literature, regional knowledge, technical expertise, and historical perspectives that may not be available online.
As one industry participant reportedly put it, the world’s best AI training data may already be sitting on bookshelves.
The result is surging demand for physical books and a growing network of brokers willing to supply them.
“Digital Preservation” or Cultural Destruction?
Some suppliers describe this process as digital preservation.
That language is convenient, but misleading.
Preservation traditionally means protecting both the information and, where culturally important, the original object.
Scanning a book before destroying it is not preservation in the fullest sense. It is extraction.
The distinction matters.
A preserved book remains available for future study. Scholars can inspect its paper, binding, annotations, ink, printing methods, and physical condition.
A destroyed book becomes a dataset.
Worse, that dataset may be locked inside a private commercial model, unavailable to libraries, historians, researchers, or the general public.
The Library of Alexandria in Warehouse Form
Libraries and archives exist because physical collections matter.
The destruction of the Library of Alexandria remains a symbol of civilizational loss thousands of years later. Its importance lies not only in the information that disappeared, but also in the arrogance of allowing accumulated knowledge to vanish.
Today’s destruction is less visible.
There is no single burning library. There are warehouses, scanning facilities, supply contracts, non-disclosure agreements, and recycling machines.
The process is decentralized, industrialized, and profit-driven.
That may make it less dramatic, but not necessarily less damaging.
Rare book dealers may receive attractive prices for old inventory, yet some have reportedly expressed private horror at watching culturally valuable materials enter a pipeline from which they will never return.
The Environmental Contradiction
The environmental irony makes the situation even more troubling.
Books require paper, ink, transportation, storage, and manufacturing resources. Destroying them adds disposal and recycling costs.
At the same time, training large AI models requires enormous amounts of electricity, water, computing hardware, and data-center infrastructure.
The race toward digital transformation is often presented as clean and futuristic. In reality, it depends on a vast physical system of servers, cooling equipment, mining, energy generation, warehouses, shipping, and waste.
Now the destruction of books must be added to that environmental equation.
Future historians may look back on this period as an age of profound shortsightedness, one in which tangible cultural heritage was sacrificed to create increasingly capable probabilistic text generators.
The Case Made by AI Defenders
Supporters of large-scale book scanning make several reasonable arguments.
AI systems may contribute to breakthroughs in medicine, science, accessibility, education, and productivity. Better training data can improve the quality and reliability of these systems.
Many books also exist in multiple copies. Digitization can make forgotten works searchable and potentially more useful.
In theory, these are strong arguments.
The problem arises when the books being destroyed are not easily replaceable.
Not every edition is identical. Not every obscure publication has been digitized. Not every regional text exists in multiple libraries.
Marginalized voices, small-language publications, local histories, community records, and unusual editions are especially vulnerable.
When the last physical copy disappears, the loss cannot be reversed.
Closed Models Do Not Equal Public Access
The argument that book destruction democratizes access also deserves scrutiny.
Digitization can expand access when the scans are placed in public archives, libraries, research databases, or open repositories.
But that is not always what happens here.
In many cases, the book’s content is ingested into a closed AI model owned by a private company. Users may be able to ask the model questions, but they cannot inspect the complete source, verify the scan, study the edition, or access the original formatting.
The knowledge becomes abstracted.
It survives only as statistical influence inside a system.
A book that once belonged to humanity’s physical record may become private infrastructure.
Non-Destructive Scanning Is Possible
There are alternatives.
Public discussion of this issue reportedly prompted Elon Musk to instruct teams at xAI to preserve rare books by scanning them without cutting off their spines.
This approach is slower and more expensive, but it respects the integrity of the original material.
Non-destructive digitization can involve:
- Overhead photography
- Cradle scanners designed for fragile books
- High-resolution imaging of covers and bindings
- Documentation of provenance and ownership marks
- Careful handling of brittle pages
- Returning books to sellers, libraries, or archives after scanning
These methods require more time, staff, and investment.
But cultural preservation has never been the cheapest option.
What Ethical AI Digitization Could Look Like
AI companies could still benefit from books without destroying literary heritage.
A responsible approach would include clear sourcing policies, rarity assessments, preservation requirements, and partnerships with libraries or universities.
Companies could prioritize common editions for destructive scanning while requiring non-destructive methods for rare, old, regional, or historically significant works.
They could also:
- Publish transparency reports about book acquisition
- Share scans with public institutions
- Preserve metadata about each physical copy
- Fund conservation programs
- Create independent review boards
- Return valuable books to archives
- Avoid anonymous bulk purchasing of culturally significant material
The goal should not be to stop digitization.
The goal should be to prevent digitization from becoming destruction.
The Need for Better Regulation
Legal rulings about fair use should not be interpreted as permission for cultural vandalism.
Copyright law asks whether copying harms the rights of an author or publisher. Heritage law asks whether an object has lasting historical or cultural value.
Those are different questions.
Policymakers could consider protections for books based on age, rarity, provenance, edition, language, or cultural significance.
Bulk buyers could be required to assess whether a book is replaceable before destroying it.
Libraries, archives, universities, and preservation experts should have a role in shaping these rules.
Without intervention, the market will continue rewarding the fastest and cheapest scanning method, even when that method causes irreversible loss.
What Individuals Can Do
Public awareness matters.
Readers, collectors, booksellers, researchers, and librarians can help protect vulnerable books by asking where unwanted collections are going and who is purchasing them.
Individuals can:
- Support independent booksellers
- Donate rare books to public institutions
- Purchase vulnerable editions for preservation
- Ask AI companies to disclose their training sources
- Support open digitization projects
- Encourage libraries to document unique holdings
- Raise concerns about destructive scanning practices
Not every old book is rare.
But rarity should be assessed before the shredder is switched on, not afterward.
Outsourcing Memory While Erasing Its Anchors
The deeper philosophical wound is this: humanity is outsourcing memory to machines while erasing the physical anchors of that memory.
Our history has always been preserved through objects.
Cave paintings, clay tablets, scrolls, illuminated manuscripts, letters, newspapers, photographs, and printed books all carry meaning beyond their literal content.
Each technological transition has created new opportunities and new losses.
But the AI transition is happening at unprecedented speed, industrial scale, and with remarkably little public visibility.
A scanned text may preserve information.
It does not necessarily preserve history.
A Choice Between Extraction and Preservation
We stand at a crossroads.
One path treats books as raw data: objects to be purchased, disassembled, scanned, consumed, and discarded once their contents have been extracted.
The other recognizes books as links in the chain of human thought: fragile, finite, and worthy of care.
The AI race shows no sign of slowing.
Without deliberate intervention, the shredders will continue working quietly, thinning the shelves of our collective past while expanding the hidden libraries inside private machines.
The books being destroyed today are not merely fuel for algorithms.
They are the voices of the dead speaking to the living.
Once those voices are separated from their physical history and their original form is destroyed, no amount of generated text can fully replace what has been lost.
We owe it to ourselves, and to those who come after us, to ensure that the pursuit of artificial intelligence does not come at the irreversible cost of our shared human legacy.