AI Firms' Book Destruction for Training Raises Concerns
Reports suggest that AI companies are purchasing and destroying rare books to train their models, a practice that has alarmed book lovers and preservationists. The destruction involves feeding book spines into wood chippers and scanning torn-out pages, which is the cheapest and fastest method for mass digitization. This approach is driven by the race among AI firms to advance their models using high-quality long-form texts found in books. The practice is feared to be more widespread than reported, potentially leading to the permanent loss of physical copies. Notably, Google patented a non-destructive book-scanning technology in 2009 that could mitigate such damage, but it is slower and more costly, and studies have found issues like page distortion and missed pages. The Internet Archive, which assists libraries in preserving aging collections, emphasizes that scanning old texts requires time and care, making it a human-intensive job. The article, sourced from Ars Technica, highlights the tension between technological advancement and cultural preservation, underscoring the need for more sustainable practices in AI training data acquisition.
Key facts
- AI companies are suspected of buying and destroying rare books to train AI models.
- The practice involves feeding book spines into wood chippers and scanning torn-out pages.
- This method is considered the cheapest and fastest way to scan books for AI training.
- Book lovers fear the practice is happening on a larger scale than reported.
- Google patented a non-destructive book-scanning technology in 2009.
- Studies have found that Google's method can distort text and miss pages.
- The Internet Archive has long understood that scanning old texts requires time and attention.
- The article was published on Ars Technica in August 2026.
Entities
Institutions
- Internet Archive
- Ars Technica