Hachette v. Google: one of the first complaints constructed to remain dangerous even if a court concludes that model training can sometimes be fair use.
Summary
The Hachette v. Google class-action complaint, filed by Hachette, Cengage, Elsevier, and Scott Turow, alleges Google unlawfully used copyrighted books and scholarly articles to train its Gemini models. The complaint claims Google acquired content from its own publishing services, web scrapes including pirate sites, and paywalled sources, then stripped copyright information and produced competing outputs. This lawsuit is strategically stronger than previous AI cases because it separates unlawful acquisition from model training, cites internal Google warnings, and provides concrete reproduction examples. While market-harm and DMCA claims need more evidence, the case is designed to remain dangerous even if model training is deemed fair use. A mixed ruling is likely, potentially finding Google liable for unlawfully sourced copies, leading to a substantial settlement involving licensing, transparency, and provenance controls.
Key takeaway
For Directors of AI/ML evaluating data sourcing for model training, this case highlights the critical risk of repurposing content obtained under limited commercial relationships or from unauthorized sources. You must meticulously audit your data acquisition channels and internal records, as courts may distinguish between transformative training and unlawful data provenance. Prioritize establishing clear licensing frameworks and robust provenance controls to mitigate substantial legal and financial exposure.
Key insights
Copyright lawsuits against AI models are strengthened by separating unlawful data acquisition from model training.
Principles
- Fair use for one function does not grant universal repurposing rights.
- Internal warnings and data provenance are critical evidence.
- Market dilution claims require robust economic analysis.
Method
The complaint separates alleged conduct into distinct stages: acquiring copies, retaining and preparing the corpus, and later model training, to address different legal vulnerabilities.
In practice
- Audit data acquisition sources and contractual terms carefully.
- Document internal risk assessments and licensing decisions.
- Implement robust provenance and attribution mechanisms.
Topics
- AI Copyright Litigation
- Generative AI Training
- Fair Use Doctrine
- Data Provenance
- Market Dilution
- Copyright Management Information
Best for: CTO, VP of Engineering/Data, Executive, Legal Professional, Director of AI/ML, AI Product Manager
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Pascal’s Substack.