AI vendors are destroying rare books to feed the chatbot
Summary
AI vendors are actively destroying rare physical books to acquire training data for chatbots, a practice exemplified by Anthropic's "Project Panama." In 2025, Judge William Alsup ruled Anthropic's destructive scanning of purchased books as "exceedingly transformative" and thus fair use, despite the company's efforts to keep the initiative secret. However, Anthropic faced a \$1.5 billion payout for training on pirated book libraries, which was not deemed fair use. The company reportedly continues to acquire rare books for destructive scanning. Bulk book buyers like Zoom Books are suspected of facilitating this process, purchasing large volumes of old books. ISBNdb, a book catalog platform, is now a key part of this chain, marketing pre-2022 print books as a "provably clean corpus" for AI training, advocating for their scanning and pulping, and assuring AI clients of anonymity. ISBNdb justifies this by stating a physical book's purpose is served once its information is encoded into an AI model.
Key takeaway
For AI Ethicists and Legal Professionals evaluating data sourcing, you must acknowledge the Anthropic ruling. This precedent permits destructive scanning of physical books as fair use for AI training. Services like ISBNdb now facilitate bulk acquisition and pulping of pre-2022 texts. This practice demands critical examination of intellectual property and cultural preservation. It also raises ethical questions about treating physical artifacts as mere data sources. Advocate for policies balancing AI innovation with safeguarding historical knowledge.
Key insights
AI vendors are systematically destroying physical books for training data, utilizing legal rulings and specialized intermediaries.
Principles
- Destructive scanning for AI training can be fair use.
- Pre-2022 print books offer "clean" training data.
- Anonymity is crucial for AI data sourcing.
Method
AI companies employ bulk buyers and catalog platforms to acquire, scan, and pulp physical books, integrating extracted text into models.
In practice
- Target pre-2022 print materials for clean data.
- Engage specialized book catalog services.
- Understand "transformative use" in data acquisition.
Topics
- AI Training Data
- Fair Use Doctrine
- Copyright Law
- Book Preservation
- Large Language Models
- Anthropic
- ISBNdb
Best for: Investor, CTO, VP of Engineering/Data, AI Ethicist, Legal Professional, Tech Journalist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Pivot to AI.