OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents

· Source: Computer Vision and Pattern Recognition · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics, Computer Vision · Depth: Expert, quick

Summary

OmniMapBench is a new benchmark designed to evaluate visual-centric reasoning in Large Vision-Language Models (LVLMs) on diverse map documents. It addresses a critical limitation in existing document understanding benchmarks where visual content is often reducible to text, allowing high performance without genuine visual grounding. Comprising 2,096 manually annotated question-answer pairs across 1,603 map documents from nine categories, OmniMapBench probes a hierarchy of skills from perception to multi-step visual reasoning. The benchmark introduces the Visual Dependency Index (VDI), a metric quantifying accuracy drop when images are replaced with text, confirming its focus on irreducible visual reasoning. Evaluations of 25 leading LVLMs revealed a significant performance gap, with the top model achieving only 75.03% accuracy, highlighting the benchmark's challenge.

Key takeaway

For AI scientists and machine learning engineers developing or evaluating Large Vision-Language Models for document understanding, you should integrate OmniMapBench into your evaluation suite. This benchmark exposes current LVLM weaknesses in genuine visual-centric reasoning on map documents, with top models reaching only 75.03% accuracy. Using OmniMapBench will help you identify and address critical gaps in your models' ability to process irreducible visual information, driving progress beyond text-reducible approaches.

Key insights

OmniMapBench provides a robust benchmark for evaluating visual-centric reasoning in LVLMs on map documents, addressing limitations of text-reducible visual content.

Principles

Method

The Visual Dependency Index (VDI) quantifies visual dependency as the accuracy drop when images are replaced with question-agnostic descriptions.

In practice

Topics

Code references

Best for: Research Scientist, AI Scientist, Machine Learning Engineer, Computer Vision Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computer Vision and Pattern Recognition.