AI companies buying books in bulk have become a new headache for booksellers and publishers, according to an investigation by 404 Media. The report says AI firms are using intermediaries to acquire large volumes of secondhand books for model training, a quieter route than scraping the web or admitting they want shelves of copyrighted human writing.
The motive is straightforward enough: the public internet is increasingly polluted with machine-generated text, the low-grade filler often called AI slop. Training a model on that material can feed the system its own recycled output. Older print books, especially those published before the recent generative AI boom, are more likely to contain human-written material that has not been mixed with AI output.
Why are AI companies buying old books?
Books are attractive training material because they contain long-form edited text, structured knowledge and writing that predates the flood of AI-generated web content. For model builders, that makes physical books a cleaner input than much of today’s internet, assuming the company can lawfully obtain and process them.
404 Media reported that ISBNdb, an online book database with more than 111 million cataloged titles, has shifted toward services that help bulk-buy books for AI companies. The reported orders range from 1,000 copies to as many as one million books in a single purchase.
An anonymous professional bookseller told 404 Media that sales began rising sharply in April. The seller said a strong week once meant moving around 20 books, while recent weekly sales reached several hundred, described as roughly five times normal volume. Other sellers on Alibris and Biblio reported similar bursts of bulk buying, according to 404 Media.
The buyers’ behavior raised flags for sellers, though 404 Media said there was no concrete proof identifying ISBNdb or any specific AI company behind every purchase. The orders reportedly focused on books with ISBNs, the 13-digit identifiers used to track editions globally. Sellers also saw little pattern by author, genre or subject, and buyers appeared willing to pay even inflated prices.
What happens to the books after scanning?
The Anthropic litigation offers the clearest public example of how physical books can enter an AI pipeline. In that case, Tom Harvey, who previously worked on Google Books and later led Anthropic’s “Project Panama” digitization effort, confirmed that Anthropic hired multiple document scanning companies, according to reporting cited in court materials.
One of those vendors was Datamation Information Services, which offers both non-destructive and destructive book scanning. Non-destructive scanning uses tools such as overhead scanners, flatbed scanners or V-shaped imaging systems. Destructive scanning removes the book binding so pages can be fed through high-speed industrial scanners. It is faster and cheaper, and it also turns the book into scrap.
Anthropic has already faced scrutiny over books. A court decision in its case found that using books to train AI could qualify as fair use, while the company faced a $1.5 billion penalty tied to a repository of seven million pirated books that infringed authors’ and publishers’ copyrights, according to prior reporting. Separately, a group of publishers has sued Google, accusing the company of unlawfully using millions of copyrighted books to develop Gemini models.
The unresolved policy problem is not subtle. If bulk buyers are pulling rare or out-of-print books out of circulation, the public loses access while private AI developers keep the resulting scans for their own systems. Better models may come from that pipeline, but the books themselves may not survive it.
This story draws on original reporting from Tom's Hardware.