Something unusual is happening in the quiet world of secondhand bookselling. Shop owners across the UK, Ireland, Australia, and the US are fielding bulk orders that follow no logical pattern. An Estonian translation of a John le Carre spy novel. A specific imprint of a Victorian classic. A 1983 issue of a niche warship magazine. Taken individually, these requests seem random. Taken together, they point to something far more deliberate and potentially transformative for the entire book trade.
A Pattern That Does Not Look Like a Pattern
Experienced booksellers know their customers. Collectors chase themes. Libraries stock by subject. Resellers buy strategically. But the buyers now flooding platforms like AbeBooks and Biblio with orders seem to follow none of these rules. One UK seller has fulfilled roughly 6,000 orders from similar buyers since January, worth thousands of pounds, with purchases spanning wildly unrelated genres and eras. Buyers are paying full asking price, sometimes routing deliveries through freight warehouses, and operating under opaque aliases.
The lack of thematic consistency is actually the tell. When a machine selects training data, it does not care whether a book is about 18th century African farming tools or 1950s motorsport biographies. It cares about the volume, diversity, and linguistic richness of the text. Diversity of content is precisely what makes a dataset valuable for training large language models. A model fed only cookbooks learns to talk about food. A model fed everything learns to talk about anything.
What Anthropic Revealed Changes Everything
The broader context here is critical. Earlier this year it emerged that Anthropic, the company behind the Claude AI assistant, spent tens of millions of dollars purchasing physical books, slicing off their spines, scanning the contents, and then recycling the physical copies. Anthropic acknowledged the practice while stating it avoids rare or antiquarian titles. That admission validated what many in the book trade had suspected and opened the door to questions about how widespread this acquisition strategy really is across the AI industry.
Books published before 2022 carry a particular appeal for AI developers because the text predates the era of AI-generated content. Training a model on human-written prose from previous decades produces cleaner, more reliable linguistic data. A database company was even caught marketing pre-2022 books directly to AI firms as ideal training material before quietly removing the promotional page. The commercial logic is clear even if the ethics remain contested.
What This Means for Buyers and the Broader Tech Landscape
For consumers watching the AI industry evolve, this story is a useful window into how aggressively technology companies are competing for the raw materials that power next-generation products. The race to build better AI is not just a software competition. It is a data acquisition arms race playing out in charity shops, online marketplaces, and independent bookstores. As buyers increasingly choose AI-powered tools and assistants based on quality and capability, understanding where that intelligence comes from matters. The hidden supply chain behind your favourite chatbot might just start on a dusty shelf in Northumberland.
