DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

· Source: Computer Vision and Pattern Recognition · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Advanced, quick

Summary

DrawingVQA is introduced as the first benchmark specifically designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings, a critical media in architecture and civil engineering. This benchmark addresses the unique complexity of construction drawings, which fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text. DrawingVQA comprises 33 "Issued for Construction" drawings and 92 expertly curated question-answer pairs, categorized into three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. A dual categorization framework is presented to analyze MLLM performance across seven construction-engineering and four MLLM capability dimensions. Initial evaluations of leading MLLMs reveal a substantial performance gap compared to human experts, particularly at higher reasoning depths, highlighting the need for advancements in domain-specialized multimodal reasoning for engineering workflows.

Key takeaway

For AI Scientists and Machine Learning Engineers developing MLLMs for architecture, engineering, and construction, you should recognize that current models exhibit a substantial performance gap on real-world construction drawings. Your development efforts must prioritize domain-specialized multimodal reasoning, particularly focusing on improving contextual interpretation and domain-expert reasoning capabilities. This benchmark indicates a critical need to advance MLLM understanding beyond perceptual tasks to integrate AI effectively into complex engineering workflows.

Key insights

MLLMs significantly underperform experts on multi-depth reasoning for complex construction drawings.

Principles

Method

DrawingVQA evaluates MLLMs using 33 "Issued for Construction" drawings and 92 Q&A pairs, categorized by three reasoning depths and a dual framework.

In practice

Topics

Best for: AI Scientist, Machine Learning Engineer, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computer Vision and Pattern Recognition.