GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]
Summary
The ActiveVision benchmark, detailed in a new arXiv paper, reveals significant limitations in frontier vision models' ability to perform "repeated visual perception." GPT-5.5, at its highest reasoning-effort tier, achieved only 10.6% on the 17 tasks across three categories, scoring zero on 11 tasks. Claude Fable 5, despite leading other reasoning and coding leaderboards, managed a mere 3.5%. In stark contrast, human participants averaged 96.1%. The critical finding is not just the low scores, but the specific nature of the failure, which models cannot circumvent by generating their own code, indicating a fundamental gap in their visual reasoning capabilities.
Key takeaway
For AI scientists and machine learning engineers developing multimodal models, recognize that current frontier vision models like GPT-5.5 and Claude Fable 5 exhibit profound weaknesses in "repeated visual perception." Your development efforts should prioritize novel architectural approaches that integrate visual reasoning more deeply, rather than relying on code generation to compensate. Consider exploring techniques like "Thinking with Visual Primitives" to address these fundamental limitations and improve model accuracy and efficiency in complex visual tasks.
Key insights
Frontier vision models struggle with repeated visual perception, a gap not solvable by code generation.
Principles
- "Language-first, vision-second" models have inherent limitations.
- Benchmarks requiring repeated visual perception expose core model weaknesses.
Method
The "Thinking with Visual Primitives" technique uses policy distillation from expert AI models to train a student model, enabling it to "point at things while thinking" for improved visual reasoning.
In practice
- Evaluate models on benchmarks like ActiveVision for "repeated visual perception."
- Explore "Thinking with Visual Primitives" for enhanced visual reasoning in AI.
Topics
- ActiveVision Benchmark
- Multimodal AI
- Visual Perception
- GPT-5.5
- DeepSeek
- Visual Reasoning
Best for: Research Scientist, AI Scientist, Machine Learning Engineer, Computer Vision Engineer
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Machine Learning.