AI Companies Are Destroying Rare Books After Using Them to Train Their Models

Artificial intelligence companies are increasingly turning to rare and antique books for training data, purchasing physical volumes, scanning every page, and in some cases dismantling the books in the process. The practice is fueling debate over whether the race to build smarter AI is coming at the expense of preserving cultural history.

Kenneth MunozJuly 29, 20262 min read

For decades, rare books have been prized for their craftsmanship, historical significance, and ability to preserve knowledge across generations. Today, some of those same volumes are finding themselves at the center of a very different mission: training artificial intelligence.

As AI developers search for new sources of high-quality text, printed books have become an increasingly valuable resource. Unlike much of the modern internet, many older works contain carefully edited language, specialized knowledge, and writing styles that can improve the diversity of training data used to develop large language models.

But getting that information into an AI system isn't always a gentle process.

Organizations involved in large-scale digitization often remove a book's spine so individual pages can pass through high-speed scanners. While the method dramatically speeds up scanning and improves image quality, it permanently destroys the physical copy being processed. For common books, the tradeoff is often considered acceptable. When older or harder-to-replace volumes are involved, the practice has sparked criticism from librarians, archivists, and historians who argue that cultural artifacts should not become disposable inputs for AI.

Supporters counter that digitization preserves the contents of a book long after the paper begins to deteriorate. Once scanned, the text can be archived, searched, and made accessible to researchers worldwide while also contributing to AI systems capable of answering increasingly complex questions.

The debate reflects a broader challenge facing the AI industry. Large language models require enormous quantities of high-quality information, yet the supply of freely available digital text is becoming more limited. That has pushed companies to explore new data sources, including books, academic publications, and historical collections that were never originally created for machine learning.

The controversy is not simply about technology. It raises fundamental questions about ownership, preservation, and the value of physical media in a digital age. Is the knowledge inside a rare book more important than the book itself? Or does destroying the original erase part of its historical significance that no digital copy can fully replace?

As artificial intelligence continues to evolve, demand for premium training data is unlikely to slow. The challenge for publishers, libraries, and AI companies will be finding ways to expand access to human knowledge without sacrificing the cultural artifacts that have safeguarded it for centuries.

The books may ultimately live on as digital records, but the growing debate suggests that how AI acquires its knowledge could become just as important as what it learns.