Secondhand booksellers across the UK and Ireland are raising questions over a series of unusual bulk purchases that they suspect could be linked to artificial intelligence companies looking for data to train their AI models.
The orders have reportedly involved large numbers of books spanning unrelated subjects, authors and genres. Instead of buying collections centred around a particular topic, some customers appear to be selecting books based on seemingly random criteria. For booksellers accustomed to dealing with collectors, libraries and individual readers, the purchasing patterns have appeared unusual enough to trigger speculation about their real purpose.
The development comes after Anthropic, the AI company behind Claude, was found to have spent millions of dollars acquiring books as part of an effort to obtain data for training its artificial intelligence systems.
The revelations have given booksellers a possible explanation for the mysterious buying activity they have been witnessing.
Unusual orders puzzle booksellers
Bulk purchases are not uncommon in the secondhand book trade. Dealers regularly sell large collections to libraries, collectors, universities and retailers. However, some of the recent orders have reportedly stood out because of the unusual combinations of titles being requested.
A buyer could ask for hundreds or thousands of books covering completely unrelated subjects, making it difficult to identify a conventional collecting strategy.
For example, an order might include literary novels, technical manuals, historical publications, biographies and specialist magazines from different decades. There may be no obvious connection between the authors, subjects or editions.

Such a purchasing strategy makes more sense if the goal is not to read the books but to obtain the information contained within them.
That possibility has become increasingly significant as AI companies search for enormous quantities of high-quality text to train increasingly sophisticated models.
Why physical books matter to AI
Artificial intelligence companies require huge datasets to train large language models. Much of the training material comes from digitally available information, including websites, public documents, books and other written material.
But a significant amount of published knowledge remains difficult to access digitally.
Millions of books have never been completely digitised. Older publications, specialist works and out-of-print titles can contain information that may not be available on the open internet.
For an AI company, acquiring physical copies of these books could therefore provide access to large amounts of unique text.
The process could involve purchasing books, scanning their pages and converting the images into machine-readable text. Once digitised, the material can potentially be processed as part of a much larger dataset.
This creates an unusual new demand for secondhand books. Titles that have little commercial value to ordinary readers could suddenly become useful because they contain text that an AI system has not previously encountered.
Anthropic’s controversial book purchases
The suspicions among booksellers have been strengthened by the case of Anthropic.
The company reportedly spent millions of dollars buying physical books for data acquisition. Its efforts became particularly controversial because some of the books were reportedly dismantled after being acquired so that their pages could be scanned more efficiently.
The approach illustrates the enormous scale of the AI industry’s demand for data.
Instead of simply purchasing digital databases, companies can seek out physical copies of books that are difficult to obtain electronically. Once the material has been digitised, the physical copies may have little further use to the company.
That has raised concerns among authors, publishers and booksellers about what happens to copyrighted material once it enters an AI company’s data pipeline.
A new source of revenue for struggling booksellers
The situation is not entirely negative for secondhand booksellers.
The industry has faced significant challenges from online retailers, changing consumer habits and the declining availability of certain physical book buyers. A sudden increase in demand for large quantities of books could provide an important source of revenue for independent dealers.
AI-related buyers could potentially purchase titles that would otherwise remain on shelves for years.
For booksellers, however, the identity and intentions of these customers matter. Some may be comfortable selling books to buyers who intend to digitise them, while others may have concerns about contributing to AI training without knowing how the material will ultimately be used.
The issue is particularly sensitive when rare or culturally significant books are involved.
Concerns over disappearing books
One of the biggest worries is that books purchased for scanning could effectively disappear from circulation.
If a company buys a rare book, scans its contents and then destroys the physical copy, the text may survive digitally, but the original object is lost.
For ordinary mass-market books, that may not be particularly significant. For rare editions, regional publications or historically important works, the situation could be very different.
Physical books can contain information beyond their printed words. Their bindings, illustrations, annotations, paper and publication history can all have cultural or historical value.
This has led to broader questions about whether AI-driven demand could unintentionally contribute to the removal of important books from the secondhand market.

Copyright questions remain
The growing use of books to train AI systems is also intensifying the debate over copyright.
Authors and publishers have increasingly questioned whether technology companies should be allowed to use copyrighted works to train AI models without obtaining explicit permission or paying creators.
AI companies, meanwhile, have argued that their use of legally obtained material can be protected under existing copyright frameworks in certain circumstances.
The legal landscape remains complicated because copyright laws differ between countries. A book purchased legally in one country could ultimately be digitised and processed elsewhere, creating additional questions about which laws apply.
The AI data race reaches bookstores
The strange bulk orders reported by booksellers highlight how far the AI industry’s demand for data has expanded.
The competition to build more capable models is no longer limited to finding powerful computer chips, constructing enormous data centres or hiring leading researchers. Companies are also searching for new sources of information that can give their models an advantage.
Secondhand bookshops may have unexpectedly become part of that race.
For now, there is no evidence that every unusual bulk purchase is connected to an AI company. Some orders could simply come from collectors, exporters, libraries or specialist dealers.
But the pattern has become difficult for booksellers to ignore.
As AI companies continue searching for new and diverse sources of training data, the humble secondhand book could become an increasingly valuable commodity. What looks like an obscure title gathering dust on a shop shelf could represent thousands of pages of information to an AI developer.
The development reflects a rapidly changing relationship between technology and traditional publishing. Books that once seemed commercially insignificant may now hold considerable value—not necessarily because someone wants to read them, but because an artificial intelligence system wants to learn from them.




