Key Takeaways
- Amazon is reportedly engaging in "destructive scanning" of physical books, including rare texts, to acquire high-quality data for training its AI models, such as Nova.
- The practice involves cutting off book bindings to speed up scanning, thereby destroying the original physical copies.
- AI companies are turning to physical books because the internet's available data has been largely exhausted, and older, human-generated texts offer superior quality, free from "AI slop."
- This controversial method raises significant ethical concerns about the destruction of cultural heritage, the privatization of knowledge, and the opaque nature of AI training data acquisition.
In a startling revelation that has sent ripples across the tech and literary worlds, Amazon, the retail giant that began its journey selling books, is now reportedly involved in a controversial practice: acquiring physical books, scanning their contents for training its artificial intelligence models, and then destroying the original texts. This news, initially brought to light by an investigation from 404 Media, highlights a growing, industry-wide trend where the hunger for high-quality training data for Large Language Models (LLMs) is clashing with ethical considerations and the preservation of cultural heritage.
The report details how a shipment containing a rare book was tracked via an AirTag to an Amazon warehouse in Las Vegas, known internally as VGT3. Employees at this facility reportedly confirm that their routine involves receiving large quantities of printed books, cutting off their bindings to facilitate faster scanning, and then discarding the now-destroyed physical copies. This process is apparently used to feed Amazon's AI models, including its Nova series.
The Scarcity of Quality Data Fuels Destructive Practices
The primary driver behind this aggressive data acquisition strategy is a looming problem for AI developers: the scarcity of high-quality training data. Large Language Models, which power conversational AI and advanced text generation, require immense volumes of text to learn from. While the internet once seemed an endless wellspring of data, much of its readily available content has already been scraped and utilized. Furthermore, the increasing prevalence of AI-generated content – often termed "AI slop" – on the internet is contaminating the data pool, making it counterproductive for training new, more sophisticated models.
This is where physical books, especially older and rare editions, become incredibly valuable. They represent a vast, untapped reservoir of human-generated language, meticulously edited, curated, and published long before the advent of generative AI. Training models on such texts can help them learn "how to write well" and develop a nuanced understanding of language, rather than merely mimicking the "low quality internet speak" that has become pervasive online.
Amazon's AI Training Facility: VGT3 and the T-Rex Logo
The 404 Media investigation provides concrete details about Amazon's operation. By embedding an Apple AirTag in a rare book, reporters were able to trace its journey to the VGT3 Amazon warehouse in Las Vegas. Here, workers confirmed the process of "destructive scanning," where book spines are cut off to allow individual pages to be fed quickly through high-speed industrial scanners. The physical book is, by design, destroyed in this process.
Adding a surreal touch to the situation, the logo for the VGT3 team responsible for this scanning operation reportedly features a Tyrannosaurus Rex holding a book in its hands, seemingly about to devour it. While Amazon has publicly stated that it "purchases books through commercial channels to help develop and improve the products and services our customers use," it has not explicitly detailed the extent or nature of this scanning operation, nor which specific AI products benefit from this data, beyond its Nova models.
A Broader Industry Trend: Anthropic and "Project Panama"
Amazon is not alone in this controversial pursuit. The practice of "destructive scanning" gained wider public attention during a much-publicized copyright lawsuit against Anthropic, the company behind the Claude AI chatbot. Court documents revealed Anthropic's "Project Panama," described as an "effort to destructively scan all the books in the world." Anthropic reportedly invested millions of dollars in this project, acquiring countless physical books, cutting their spines, digitizing their contents, and then disposing of the originals.
While Anthropic claimed in a statement to the BBC that "none of our data acquisition programs buy and destroy rare or antiquarian books," the revelations from their lawsuit and the recent Amazon investigation suggest that the demand for diverse, high-quality textual data is pushing AI companies to explore all available avenues. Indeed, ongoing lawsuits against other tech giants like Meta, Google, Microsoft, and OpenAI hint at similar practices across the industry, with a particular focus on works by deceased authors who cannot protest.
Ethical and Legal Crossroads: Fair Use vs. Cultural Destruction
The practice of destructive scanning raises profound ethical and legal questions. From a legal standpoint, a federal judge in the Anthropic case ruled that using legally purchased and scanned books to train an AI model could qualify as "fair use" under copyright law. This ruling was partly based on the argument that the physical originals were destroyed and not copied for resale, making the digital use "transformative." However, this "fair use" defense did not protect Anthropic from a staggering $1.5 billion settlement for maintaining a repository of millions of pirated books that infringed copyrights. Similarly, a coalition of publishers has sued Google, alleging illegal use of copyrighted books for its Gemini AI models.
Beyond copyright, the ethical implications are far-reaching. The destruction of rare and out-of-print books means the irreversible loss of physical artifacts that carry historical, cultural, and aesthetic value far beyond their mere textual content. Annotations, unique bindings, inscriptions, and the provenance of a specific copy contribute to its cultural significance, all of which are lost in destructive scanning. This process effectively privatizes knowledge, pulling irreplaceable information off public shelves and locking it within the proprietary, closed AI models of a single corporation.
Libraries, archivists, collectors, and booksellers are increasingly concerned. The opacity of these acquisition processes, often conducted through anonymous purchases on marketplaces, makes it difficult to track what is being bought and destroyed. While some libraries are partnering with AI companies for ethical digitization projects, aimed at open access and preservation, the current trend of destructive scanning represents a stark contrast to these efforts.
The Future of Knowledge and Preservation in the AI Era
The debate around AI training data highlights a critical juncture for how humanity manages and preserves its collective knowledge. On one hand, AI offers powerful tools for digital preservation, making fragile manuscripts accessible and searchable through advanced imaging, OCR, and transcription technologies. Libraries, such as the Library of Congress, are actively engaging with AI companies to explore ethical partnerships for digitizing vast archives.
However, the aggressive, destructive methods employed by some AI companies threaten to undermine these preservation efforts. The rush to acquire data risks prioritizing immediate AI development over long-term cultural stewardship. International regulations, such as the EU AI Act, are beginning to mandate transparency in AI training data, requiring providers to disclose summaries of their data, including sources and licensing. Such regulations could be crucial in fostering more responsible data acquisition practices.
Ultimately, the challenge lies in finding a balance: enabling AI to learn from the rich tapestry of human knowledge without destroying the very sources that embody it. As AI models become more integrated into our lives, the decisions made today about their training data will profoundly shape the accessibility, integrity, and future of human knowledge for generations to come.
Frequently Asked Questions
What is "destructive scanning" in the context of AI training?
Destructive scanning is a process where physical books are disassembled, typically by cutting off their spines, to allow individual pages to be fed quickly through high-speed industrial scanners. This method is efficient for digitizing large volumes of text but results in the irreversible destruction of the original physical book.
Why are AI companies, like Amazon, interested in scanning physical books, especially rare ones?
AI companies are increasingly turning to physical books because the vast majority of online text data has already been used for training Large Language Models (LLMs). Older, rare, and out-of-print books contain high-quality, human-generated language that is not readily available online and predates the proliferation of AI-generated content, making it invaluable for improving AI model performance and teaching them to "write well."
What are the ethical concerns surrounding the destruction of rare books for AI training?
The practice raises significant ethical concerns, including the irreversible loss of cultural heritage (as the value of a book extends beyond its text to its physical form, annotations, and history), the privatization of knowledge (locking public information within corporate AI models), and the opaque nature of data acquisition.
Is it legal for AI companies to scan and destroy copyrighted books for training?
In the U.S., a federal judge in a lawsuit against Anthropic ruled that "destructive scanning" of legally purchased books for AI training could qualify as "fair use" under copyright law, partly because the originals were destroyed and not resold. However, using pirated or illegally obtained copyrighted material for training has led to significant fines and legal battles for AI companies.



