Claims that AI firms destroying rare books are already a fact run ahead of the evidence. Booksellers in Europe have reported odd, high-volume orders that they suspect are headed for AI training pipelines. Destructive scanning of purchased print books has been documented in Anthropic's case. What has not been established is that the anonymous buyers are AI companies, or that rare and collectible copies are being cut apart after sale.
That distinction is doing a lot of work. A book can be cheap as reading material and irreplaceable as an object, carrying a particular binding, notes in the margins, a signature, a dedication, a bookplate, or features unique to an edition.
Are AI companies destroying rare books?
There is no confirmed evidence in the available reporting that AI companies are destroying rare books. Ars Technica reported that a 2025 lawsuit revealed Anthropic had destructively scanned millions of legally purchased print books for model training, but said there was no indication the company had destroyed rare books. Anthropic has denied doing so, according to Ars Technica.
The concern is not invented from thin air. Chosun reported accounts from secondhand and antique shops of bulk orders that did not resemble normal institutional collecting. One Galway, Ireland, bookseller reportedly received an order for 5,000 books spanning local-interest titles and an old investment guide. Other sellers described eclectic lists, no requests for volume discounts, ISBN-driven selection patterns, and delivery through logistics hubs before onward shipment to the United States.
Those patterns suggest automated acquisition rather than a reader building a library. They do not identify the purchaser. Chosun said it remained unconfirmed whether particular books sold through those orders reached named AI companies or what happened to them afterward.
Why old print books appeal to AI-data buyers
ISBNdb, which 404 Media reported offers high-volume acquisition of print books for AI companies, markets older books as material produced before widespread generative-AI publishing. The company calls books curated and domain-specific human knowledge, and contrasts them with web data. That is ISBNdb's sales pitch, not an independent measurement of training-data quality.
For model developers, books can offer long-form, edited text that is less likely to have passed through a chatbot before reaching the scanner. For preservationists, the same market creates a problem: a buyer seeking text may treat a copy's physical history as disposable.
Destructive scanning is fast. Preservation scanning is slower.
Destructive scanning removes a book's binding, separates its pages and feeds them through high-speed equipment. Ars Technica reported that this is the cheaper and faster route for scanning at scale. It also ends the life of that particular copy.
The Internet Archive uses a different approach for fragile and rare material. Its description of the process, which the organization told Ars Technica still reflects its current practice, relies on a human operator turning pages, checking images and returning to foldouts or missed material. The Archive said automated vacuum page-turners had not worked well for brittle books, rare volumes and special collections.
OpenAI and Microsoft have also worked with Harvard librarians on a project involving roughly 1 million public-domain books dating to the 15th century, according to Ars Technica. That is a digitization effort, not evidence about the unidentified bulk buyers now worrying sellers.
The available record supports a narrower conclusion than the viral version: destructive book scanning for AI training has happened, while claims that rare books are being systematically bought and destroyed remain a plausible preservation risk, not a confirmed supply chain.
This story draws on original reporting from Ars Technica.