Amazon is destroying rare books to train AI

📡 TechCrunch · 1 min read ·
Rare books are becoming a key resource for training large language models, or LLMs—the technology behind tools like ChatGPT. These models have already absorbed most of what is freely available online, so companies are now turning to physical texts that have never been digitized. Amazon, which began as an online bookstore, is reportedly destroying some of these rare volumes to extract their content for AI training. The process involves scanning or cutting apart the books to capture text that is not accessible elsewhere. This practice raises concerns among collectors and researchers, who see the destruction as a loss of cultural heritage. However, the demand for unique data is growing as AI developers seek to improve the accuracy and depth of their models. The move highlights a broader shift: as online data becomes exhausted, the value of physical archives is rising—even if that means sacrificing them in the process.