AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
Summary
AI companies are increasingly purchasing vast quantities of old, printed books, primarily those published before 2022, to use as training data for their large language models. This trend is driven by the need for "clean" data, free from AI-generated text that can lead to "model collapse." ISBNdb, a company specializing in book metadata, now facilitates bulk acquisitions of 1,000 to 1 million books for AI labs, emphasizing that print books offer "curated, peer-reviewed, domain-specific human knowledge." This practice often involves "destructive book scanning," where books are destroyed, leading to "optics problems" and secrecy measures by ISBNdb. Despite copyright concerns, a federal judge ruled Anthropic's destructive scanning for internal digital copies as "transformative" fair use. Booksellers report an unprecedented surge in sales of diverse, often rare, and foreign-language books, indicating significant demand from AI entities.
Key takeaway
For Directors of AI/ML seeking high-quality training data, you should prioritize sourcing pre-2022 print materials to mitigate "model collapse" risks from AI-generated content. Be aware that bulk acquisition often involves destructive scanning, which, while legally deemed fair use for internal copies, carries significant "optics problems." You must weigh the benefits of clean data against potential public perception issues and consider implementing robust NDAs for your sourcing operations.
Key insights
Pre-2022 print books are a critical source of "clean" AI training data, free from model-degrading AI-generated content.
Principles
- Data quality is paramount for AI model stability.
- Destructive copying for internal use can qualify as fair use.
- Secrecy is valued to manage public perception.
Method
AI companies acquire bulk print books via services like ISBNdb, often using destructive scanning for efficient digitization, then process them into training datasets.
In practice
- Source pre-2022 print materials for uncontaminated datasets.
- Implement NDAs for sensitive data acquisition strategies.
- Consider destructive scanning for cost-effective digitization.
Topics
- AI Training Data
- Model Collapse
- Destructive Scanning
- Copyright Law
- Fair Use
- ISBNdb
Best for: CTO, VP of Engineering/Data, AI Architect, Tech Journalist, Legal Professional, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by 404media Feed.