Tue 21 Jul 2026 / 13:56 ET
Kernel
Security 4 min read

AI labs are chasing print books as cleaner training data

ISBNdb is pitching bulk book buying to AI companies that want pre-LLM text, while sellers report strange spikes in ISBN-driven orders.

Dana Voss

By Dana Voss / Security Correspondent

AI labs are chasing print books as cleaner training data
img: 404 Media

AI companies looking for more training material are being sold a very old-fashioned answer: buy printed books, scan them, and avoid the synthetic text now flooding the web.

ISBNdb, a company that says it runs the world’s largest book database, is marketing bulk book acquisition to AI labs. On its website, the company argues that books contain edited, structured human writing that web crawls cannot reliably match. Its pitch is blunt: pre-2022 print books are useful because they were published before large language models made AI-generated text cheap and common.

That matters for model builders because training data is getting messier. ISBNdb says modern web data may include AI-written material, which can contribute to “model collapse,” a term used for degradation that can occur when models train on outputs from other models. The company also warns that authors who object to scraping can deliberately write material meant to corrupt future models.

ISBNdb’s answer is the ISBN, the commercial identifier printed on most books. The company has long served booksellers, libraries and distributors with book metadata. It now says it can help AI companies source orders ranging from 1,000 to 1 million printed books, using its data to avoid buying duplicates and to target specific titles.

The company also advertises confidentiality. ISBNdb says each engagement begins with a non-disclosure agreement and that a customer’s identity, strategy and acquisition targets are not revealed. Its site acknowledges the public relations problem directly, saying a headline about an AI company destroying millions of books would not win sympathy.

The practice drew broader attention after authors suing Anthropic over copyright claims obtained internal documents describing the company’s plan to buy and scan millions of physical books. The Washington Post reported that Anthropic bought books from Better World Books, a marketplace used by libraries, retailers and individuals. Google has also been sued by publishers over claims that it trained Gemini on copyrighted books.

404 Media reported that one professional bookseller specializing in foreign-language books has seen sales jump sharply since April. The seller, who was not named so he could continue using the marketplaces, said he used to sell about 20 books in a good week and has since been selling hundreds. He said he did not have proof the buyers were AI companies, but the order patterns looked unusual: large purchases across unrelated subjects, apparently tied by the presence of ISBNs, while rare books without ISBNs were left alone.

The bookseller said the sales help him financially and clear slow-moving stock, but he dislikes the suspected end use and worries that uncommon books could be destroyed. Other booksellers have raised similar questions. On an Alibris seller forum in February, one seller asked about an increase in automated buy orders and wondered whether AI was involved. Mike Feldman, Alibris’s director of client services, replied that new bulk buyers were purchasing trade books from many sellers.

Another bookseller told 404 Media that large orders were also appearing through Biblio, where customers can submit spreadsheets of ISBNs for purchase. NL Times reported in June that rare-book dealers in the Netherlands had seen similar bulk buying and suspected tech firms.

The buyer identities remain hard to confirm because marketplaces and sourcing firms keep customers hidden. Large orders may be routed through distribution centers before reaching the final client.

Court records in the Anthropic case describe one scanning contractor, Datamation, which offers both destructive and non-destructive book scanning. In destructive scanning, the spine is cut off so pages can be fed through machines faster and at lower cost.

U.S. District Judge William Alsup ruled that Anthropic’s creation of digital copies from purchased print books was fair use, emphasizing that the originals were destroyed and the digital copies were not shown, shared or sold outside the company. ISBNdb cites a similar argument on its website, saying secondhand book purchases do not deprive creators of income from books already sold.

ISBNdb and Anthropic did not respond to requests for comment, according to 404 Media.

This story draws on original reporting from 404 Media.

More Security/

view all ↗