State of Data — Sean Cai, Independent / State of Data
Summary
The data market is undergoing significant unbundling and fragmentation, moving away from vertically integrated giants like Scale AI, Remarque, and Surge towards specialized vendors. This shift is driven by the realization that quality does not scale linearly with quantity, leading labs to diversify vendors (20-30 different ones). The focus is now on "Type One" process-based data—pure captures of real workflows like GitHub commits—rather than "Type Two" contrived data, which is often sold as Type One. Model improvement, a function of compute, data, and talent, sees data as the underfunded leg, crucial for transforming generalist models into experts. The industry struggles with "benchmark psychosis," where many benchmarks are quietly fake, testing only isolated questions and exhibiting high false positive/negative rates, as seen in comparisons between Opus 4.8 and 4.7 or GPT 5.5 and Opus 4.08 on finance tasks. Successful data companies are pivoting to enterprise solutions, building abstraction layers and "Antikythera mechanisms" to manage RL datasets and enable automatic post-training on new open-source models.
Key takeaway
For Directors of AI/ML evaluating data acquisition strategies, recognize that relying solely on vendor-provided benchmarks or contrived datasets is a significant risk. Your team should prioritize sourcing Type One, process-based data from real-world workflows and implement robust, agnostic benchmarking with detailed rubric analysis to accurately assess model performance and guide post-training efforts. This approach mitigates "benchmark psychosis" and builds a durable pipeline for continuous model improvement.
Key insights
Data markets are fragmenting, shifting to real-world process data, and current benchmarks are often unreliable.
Principles
- Data supply chains unbundle as markets mature.
- Quality does not scale linearly with data quantity.
- Verifier's Law: task verifiability dictates model training ease.
Method
Predict application layer maturity by classifying professional tasks on three verifiability axes (asymmetry, veracity, proliferation) and analyzing raw data for five specific signals like sequential decisions and inferable expert actions.
In practice
- Diversify data vendors to ensure quality.
- Prioritize Type One, process-based data for realism.
- Implement agnostic benchmarking with rubric analysis.
Topics
- Data Markets
- AI Data Supply Chain
- Benchmark Psychosis
- Verifier's Law
- Type One Data
- Enterprise AI
Best for: Research Scientist, Investor, CTO, AI Scientist, Machine Learning Engineer, Director of AI/ML
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by AI Engineer.