Companies are purchasing physical books because books published before the widespread adoption of generative AI offer something...
Summary
AI companies are actively acquiring and sometimes destroying physical books published before 2022 to secure high-quality, human-written training data untainted by generative AI output. This practice, facilitated by services like ISBNdb which helps clients acquire 1,000 to one million books per order, risks removing rare works from circulation and enclosing cultural knowledge within private datasets. The destruction, often for faster scanning, raises concerns about the loss of original artifacts and the transfer of publicly accessible material into opaque corporate systems. This trend highlights a growing scarcity of "clean" human data, potentially leading to "model collapse" from recursive training on synthetic content, higher costs for model improvement, and a concentration of AI development among firms with proprietary data access.
Key takeaway
For Directors of AI/ML evaluating data acquisition strategies, recognize that relying solely on bulk-purchased, destructively scanned books introduces significant ethical and performance risks. Your teams should prioritize transparent, licensed data sources with clear provenance to avoid "model collapse" and ensure long-term sustainability. Invest in non-destructive digitisation and quality assurance to preserve cultural assets, maintain public trust, and secure high-quality human knowledge.
Key insights
AI's reliance on human-generated data is creating a scarcity crisis, leading to unsustainable practices like destroying old books.
Principles
- Recursive training on synthetic data causes "model collapse."
- Human-created knowledge is AI's scarce natural resource.
- Data provenance and quality are critical for model performance.
Method
ISBNdb offers large-scale book acquisition services, using ISBNs to locate titles, avoid duplication, and organize destructive scanning for AI training.
In practice
- Implement transparent, non-destructive digitisation practices.
- License content directly from authors and publishers.
- Establish robust OCR quality control for scanned texts.
Topics
- AI Training Data
- Model Collapse
- Data Scarcity
- Generative AI Contamination
- Book Digitisation
- Data Ethics
- Cultural Preservation
Best for: Research Scientist, Investor, CTO, AI Scientist, AI Ethicist, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Pascal’s Substack.