Across several countries, second-hand booksellers have reported an unusual surge in bulk orders of predominantly non-fiction titles from a US-based buyer identified as Zoom Books. These orders, often comprising hundreds of diverse subjects ranging from local history to transport and geology, have raised concerns that the purchases are intended for destructive scanning to create training datasets for artificial intelligence (AI) models.
Books published before 2022 are prized in this context as they are less likely to contain AI-generated content. Experts caution that training AI systems on texts produced by other AI can cause "model collapse," a feedback loop where the quality of generated content deteriorates over time. Yarin Gal, Associate Professor of Computer Learning at Oxford University, describes this phenomenon as an “AI echo chamber” that moves further away from reliable human knowledge. Therefore, high-quality, human-written texts, especially older works, are considered essential for developing language models.
Booksellers in the UK, Germany, New Zealand, and Australia have described receiving such orders, often noting the arbitrary combinations and obscure nature of the titles requested. Finn Winde, a bookseller in York, recounted an order of mixed academic and general-interest books sent to Zoom Books, which he and his colleagues presume to be operating on behalf of a larger AI firm. Similarly, Helen Bott of Treasure Chest Books in Felixstowe highlighted the scale and randomness of the orders and expressed ambivalence—while the business is financially beneficial, the destruction of books, some of which are out of print and rare, is troubling.
The Booksellers Association has voiced concern over this emerging practice. Meryl Halls, the association’s managing director, called it a worrying example of technology companies potentially exploiting copyright-protected material without sufficient transparency or royalty compensation. Conversely, not all booksellers share this apprehension. Patrick Kelly of Book Mongers in Brixton viewed the orders pragmatically, noting that smaller shops without a detailed inventory system are unlikely to attract such transactions and expressed a level of acceptance of the business opportunity.
Representatives from AI companies have responded to the reports with some clarifications. A spokesperson for Anthropic, developer of the Claude language model, stated that its training data includes publicly available web content, commercially acquired datasets, and internally generated data, and denied involvement in buying or destroying rare or antiquarian books. Similarly, an Amazon spokesperson confirmed that books are purchased through standard commercial channels to support service development.
The practice has drawn comparisons to historical instances of book destruction, such as the infamous “Bonfire of the Vanities” in Renaissance Florence. However, in the current scenario, the motivation is linked to the breadth and depth of AI training rather than ideological or religious censorship. As AI technology companies accumulate significant wealth and influence, the bulk purchasing and destruction of second-hand books highlight tensions between commercial interests, intellectual property rights, and cultural preservation.
