Chartography: A Benchmark for Professional Chart Understanding
Summary
Chartography is a new benchmark designed to measure frontier models' ability to interpret complex, domain-specific charts with professional accuracy. Developed with tasks written and reviewed by domain experts, it includes formats like Sankey diagrams, candlestick charts, contour maps, and Kaplan-Meier curves from fields such as medicine, engineering, and finance. Unlike existing benchmarks where models score 80-90%, Chartography reveals that the best frontier model, GPT-5.6 Sol (Max reasoning), achieves only 45%, with most models scoring 10-40%. The benchmark highlights model failures in visual element detection, value estimation, complex geometry interpretation, applying domain conventions, and error propagation in multi-step reasoning.
Key takeaway
For AI Product Managers developing professional agents, this benchmark reveals a critical gap: current frontier models cannot reliably interpret complex, domain-specific charts. You should prioritize robust visual reasoning capabilities beyond simple data extraction, focusing on multi-step inference, geometric interpretation, and adherence to professional conventions. Your development efforts must address these limitations to transition from impressive demos to dependable, deployable systems in fields like medicine or engineering.
Key insights
Frontier models struggle significantly with professional-grade chart interpretation, achieving only 45% on a new expert-designed benchmark.
Principles
- Charts from real professional domains are crucial for robust evaluation.
- Questions by real professionals reveal expert-level reasoning gaps.
- Visual reasoning must go beyond simple label lookup.
Method
Chartography tasks, authored and reviewed by domain experts, require models to estimate unlabeled values, trace features, interpret geometry, apply conventions, and combine visual readings. Expert grading provides answers, walkthroughs, and chart-specific acceptable ranges.
In practice
- Evaluate models on Kaplan-Meier curves and Bode plots.
- Test interpolation between contours and curves.
- Assess understanding of domain-specific conventions.
Topics
- Chart Understanding
- Benchmark
- Multimodal AI
- Visual Reasoning
- Domain-Specific Charts
- Model Evaluation
- GPT-5.6 Sol
Code references
Best for: Research Scientist, Computer Vision Engineer, AI Scientist, Machine Learning Engineer, AI Product Manager
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Surge AI Blog.