Burning Books

AI Firms Destroying Millions of Rare Books for Training Data, Raising Alarms Over Cultural Heritage

  • AI companies are purchasing millions of physical books from secondhand markets, destroying them after scanning for training data.
  • Federal Judge William Alsup ruled Anthropic’s destructive scanning of legally purchased books qualifies as “fair use” under copyright law.
  • Pre-2022 printed books are sought because they contain no AI-generated text, avoiding “model collapse” during training.
  • ISBNdb and other intermediaries facilitate bulk purchases of 1,000 to 1 million books per order while keeping AI buyers anonymous.
  • Rare, foreign-language and out-of-print volumes face potential extinction as the last physical copies are destroyed.

(Natural News)

The judgment that changed everything

On July 21, 2026, U.S. District Judge William Alsup issued a landmark ruling in Anthropic v. Authors Guild that has fundamentally altered the relationship between artificial intelligence companies and humanity’s printed heritage. The judge declared that Anthropic’s practice of purchasing physical books, slicing off their spines for industrial scanning, and destroying the originals constitutes “fair use” under Section 107 of the Copyright Act. This decision, the first of its kind applying copyright law to AI training datasets, has unleashed an unprecedented race among technology companies to acquire and destroy books at industrial scale — threatening rare volumes and erasing physical copies of human knowledge that may never be replaced.

ADVERTISEMENT

Why pre-2022 books became gold

The driving force behind this destruction is what researchers call “model collapse.” When AI systems train on text generated by other AI systems, quality degrades with each generation, producing increasingly incoherent results. The internet has become so saturated with synthetic content that companies now actively seek pre-2022 printed books as the last uncontaminated reservoir of human-authored knowledge.

ISBNdb, a company that sources printed books for AI training data, advertises on its website that “the world’s best AI training data is sitting on a shelf.” The company describes printed books as “curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate.” Pre-2022 publications are “structurally guaranteed to be free of this contamination,” the company states, referencing both AI-generated text and the growing practice of authors “poisoning” web content to sabotage scraping pipelines.

The industrial process of destruction

The mechanics of this operation are precise and troubling. Standard pallets of books hold 800 to 1,200 volumes. Buyers scale from pilot orders to 10,000 or more books per batch. Destructive scanners process 80 to 120 pages per minute after hydraulic cutting machines slice off book spines. The original books are then pulped and recycled.

According to court documents from the Anthropic case, the company’s “Project Panama” spent tens of millions of dollars on this exact pipeline, contracting with Datamation for scanning services. Booksellers have identified these bulk buyers through telltale signs: abnormal volume, subject-agnostic orders spanning history, botany, regional law, and German economics simultaneously, and total indifference to pricing. As one bookseller told 404 Media, “It’s not just the quantity, but the weirdness of the orders.”

Anthropic’s Tom Harvey, who previously helped create Google Books, oversaw the operation. One bookseller who has sold hundreds of books to suspected AI buyers expressed mixed feelings: “It benefits me financially… On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”

The vanishing heritage

The most troubling aspect of this practice is what it means for rare and out-of-print books. Foreign-language volumes, low-circulation academic works, and antique texts with limited surviving copies face potential extinction. When an AI company purchases the last known physical copy of a rare book, scans it, and destroys the original, that volume ceases to exist in the physical world. It becomes data locked inside corporate servers — inaccessible to future scholars, collectors, or the public.

ISBNdb acknowledges the “optics problem” on its website, noting that “‘AI company destroys two million books’ is not a headline that generates sympathy.” The company promises strict nondisclosure agreements on every engagement, stating that “your identity, strategy, and acquisition targets are never disclosed.” This secrecy means booksellers can only speculate about who is buying their inventory and for what purpose.

Rare booksellers in the Netherlands have reported similar suspicious bulk purchases. One seller on Alibris forums asked in February 2026: “Is an AI going to read every single book? Any insight into how the selections are made?”

A legal framework that encourages destruction

Judge Alsup’s ruling explicitly validated the buy-scan-destroy pipeline, writing that “every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other.” Because the digital copy was never shown, shared, or sold outside the company, the judge found this “clearly transformative” and therefore protected by fair use.

This reasoning effectively creates a legal incentive for destruction. If a company keeps the physical book, it must store it. By destroying the original, the company can argue it “replaced” the physical copy with a digital one. The ruling handed the entire AI industry a template, and companies now explicitly cite it.

Harvard University, partnering with Google and Microsoft, demonstrated that an alternative exists. The collaboration released nearly one million public-domain digitized books in 254 languages — no destruction required. Microsoft’s Burton Davis called starting with public-domain data “prudent,” noting that libraries hold “significant amounts of interesting cultural, historical and language data.” That alternative track makes clear that labs have cleaner options. They simply choose not to take them.

Who controls humanity’s textual heritage?

The question of who controls humanity’s textual heritage can no longer be deferred. Every out-of-print volume that disappears into an industrial scanner may represent the last copy accessible to anyone outside an AI laboratory’s servers. Rare books, foreign-language texts, and obscure academic works — many never digitized elsewhere — vanish from physical existence after a single scan. The court’s reasoning that destruction is “transformative” has unleashed a systematic erasure of physical archives, not for preservation but for corporate profit. As AI companies accelerate their acquisition of printed knowledge, society must confront whether the convenience of training data justifies the permanent loss of books that cannot be replaced.

Sources for this article include: