ISBNdb: Save the Archive and Stop Feeding Rare and Out-of-Print Books Into the Shredder

421

Let’s get to 500 signatures!
Petitions with 1,000+ supporters are 5x more likely to win!

The Issue

ISBNdb, built on a database of more than 111 million books, now brokers bulk purchases of physical books for AI companies. Those books are shipped by the pallet, cut from their bindings, run through high-speed scanners, and pulped.

We are asking ISBNdb to stop brokering the sale of rare, scarce, and out-of-print books into destructive-scanning pipelines, and to *require* non-destruction terms on every bulk sale it facilitates.

In June 2025, court documents in Bartz v. Anthropic revealed that Anthropic had spent many millions of dollars buying used print books, stripping the bindings, slicing the pages, scanning them, and discarding the originals. As of *last year*, millions of volumes had been destroyed. Judge William Alsup ruled the practice fair use, reasoning that a legally purchased book converted to digital form and then destroyed creates no new copy. Anthropic separately settled claims over pirated digital books for $1.5 billion.

So AI wins in the court room. And they'll destroy the archive unless we get involved. Because nobody else asked.

Reporting by 404 Media in July 2026 documented that ISBNdb was marketing printed books to AI firms precisely because they predate the flood of machine-generated text: books published before large language models are guaranteed to be human-written, edited, and reviewed. ISBNdb's own pitch describes the world's best AI training data as sitting on a shelf. Booksellers have described receiving orders for hundreds of volumes at a time (regional history, botany, foreign-language economics, obscure legal texts) with nothing in common except that each title carries an ISBN. ISBNdb has acknowledged to clients that the optics are a problem.

But the real problem is the destruction of the physical archive. Many of these books exist in only a few hundred copies worldwide. Some exist in a few dozen. A regional press monograph, a self-published local history, a translated technical manual from a defunct publisher: once the surviving copies are cut apart, the object is gone permanently, and what remains is a proprietary scan inside a private corpus that no scholar, librarian, or member of the public can consult. Even if you stumble on the information in an AI search, it will have been garbled and summarized according to whatever prediction metrics the bot has been trained for. No second opinions or verifications.

Sanitized summaries (at best) are not a book. They do not preserve marginalia, provenance, bindings, paper stock, printing variants, inserted ephemera, or the physical evidence that book historians, bibliographers, and conservators depend on. It does not preserve anything at all if the company holding it is acquired, sued, restructured, or simply decides the file is no longer worth storing. Digital surrogates held privately are not an archive. They are a corporate asset with a shelf life.

None of this is necessary. The Internet Archive pioneered non-destructive book scanning at scale. Google Books digitized millions of volumes with a patented camera process and returned them to the libraries that lent them. OpenAI and Microsoft partnered with Harvard's libraries to train on nearly a million public domain books that are fully digitized and fully preserved.

Destructive scanning is chosen because it is faster and cheaper and it *forces* you to rely on AI for the final word.

Here is what we want:

  • ISBNdb should refuse to broker rare, scarce, and out-of-print titles into destructive pipelines. 
  • screen titles against WorldCat/OCLC holdings and comparable registries before fulfilling any bulk order for AI training use. Any title below a defined scarcity threshold is not for sale into this market.
  • Require a written non-destruction commitment on every bulk sale to AI purchasers.
  • Scan books non-destructively and retain, resell, or donate them to a library, archive, or Friends of the Library sale.
  • Publish transparency reporting. Disclose the annual volume of books sold for AI training, the categories and publication-date ranges involved, and the scarcity screening applied.
  • Direct AI clients toward preservation partners. Route buyers to the Internet Archive, HathiTrust, and university library digitization programs, which can supply high-quality text without pulping the source.
  • Fund what you profit from. Commit a percentage of AI acquisition revenue to non-destructive digitization at public and academic libraries.

The AI companies are culpable, too. We should demand they adopt non-destructive scanning as standard practice. Deposit every volume they acquire with a library or archive after digitizing it. Publish their acquisition policies.

Sign this petition. We are asking one company positioned at the center of this trade to draw a line, and we are asking the companies buying from it to stop pretending there was no other way.

Signed by readers, librarians, archivists, booksellers, teachers, scholars, and anyone who thinks the physical record is worth more than a marginal gain in training data.

The Decision Makers

ai
ai

Supporter Voices

Petition Updates