Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Summary
A new benchmark evaluated six Multimodal Large Language Models (MLLMs) on scientific visualization (SciVis) literacy, using a standardized assessment of 49 items across 18 scientific visualizations, 8 techniques, and 11 task types. Comparing three closed-source and three open-source models against 485 human participants, the study found MLLMs lack uniform SciVis literacy. Gemini emerged as the strongest model, surpassing the human mean, while open-source models performed below human baseline. Models excelled in scientific illustration, search, and spatial understanding tasks but struggled significantly with texture-based and integration-based visualizations, and quantitative estimation. Error analysis revealed consistent failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings underscore SciVis literacy as a critical, necessary benchmark dimension for multimodal AI system evaluation.
Key takeaway
For Machine Learning Engineers deploying MLLMs for scientific visualization interpretation, understand that current models exhibit highly uneven literacy. While Gemini shows promise, exceeding human performance, open-source alternatives generally fall short. You should carefully select models based on task type, utilizing MLLMs for spatial understanding or scientific illustration, but exercising caution and implementing human oversight for tasks requiring fine-grained quantitative estimation or interpreting complex texture-based visualizations.
Key insights
MLLMs show uneven scientific visualization literacy, with Gemini outperforming humans while open-source models lag.
Principles
- SciVis literacy is a critical MLLM benchmark.
- MLLM performance varies significantly by SciVis type.
- Quantitative estimation remains a major MLLM weakness.
Method
Benchmarked 6 MLLMs on a 49-item SciVis literacy test, comparing against 485 human participants under a closed-world protocol.
In practice
- Prioritize Gemini for SciVis interpretation tasks.
- Avoid MLLMs for precise quantitative SciVis analysis.
- Focus MLLM development on texture/integration-based SciVis.
Topics
- Multimodal LLMs
- Scientific Visualization
- AI Benchmarking
- Visualization Literacy
- Gemini
- Quantitative Estimation
Code references
Best for: AI Engineer, AI Scientist, Machine Learning Engineer, Research Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.