State of Data — Sean Cai, Independent / State of Data

· Source: AI Engineer · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics · Depth: Expert, long

Summary

The data market is undergoing significant unbundling and fragmentation, moving away from vertically integrated giants like Scale AI, Remarque, and Surge towards specialized vendors. This shift is driven by the realization that quality does not scale linearly with quantity, leading labs to diversify vendors (20-30 different ones). The focus is now on "Type One" process-based data—pure captures of real workflows like GitHub commits—rather than "Type Two" contrived data, which is often sold as Type One. Model improvement, a function of compute, data, and talent, sees data as the underfunded leg, crucial for transforming generalist models into experts. The industry struggles with "benchmark psychosis," where many benchmarks are quietly fake, testing only isolated questions and exhibiting high false positive/negative rates, as seen in comparisons between Opus 4.8 and 4.7 or GPT 5.5 and Opus 4.08 on finance tasks. Successful data companies are pivoting to enterprise solutions, building abstraction layers and "Antikythera mechanisms" to manage RL datasets and enable automatic post-training on new open-source models.

Key takeaway

For Directors of AI/ML evaluating data acquisition strategies, recognize that relying solely on vendor-provided benchmarks or contrived datasets is a significant risk. Your team should prioritize sourcing Type One, process-based data from real-world workflows and implement robust, agnostic benchmarking with detailed rubric analysis to accurately assess model performance and guide post-training efforts. This approach mitigates "benchmark psychosis" and builds a durable pipeline for continuous model improvement.

Key insights

Data markets are fragmenting, shifting to real-world process data, and current benchmarks are often unreliable.

Principles

Method

Predict application layer maturity by classifying professional tasks on three verifiability axes (asymmetry, veracity, proliferation) and analyzing raw data for five specific signals like sequential decisions and inferable expert actions.

In practice

Topics

Best for: Research Scientist, Investor, CTO, AI Scientist, Machine Learning Engineer, Director of AI/ML

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by AI Engineer.