Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics, Human-Computer Interaction · Depth: Advanced, quick

Summary

A new benchmark evaluated six Multimodal Large Language Models (MLLMs) on scientific visualization (SciVis) literacy, using a standardized assessment of 49 items across 18 scientific visualizations, 8 techniques, and 11 task types. Comparing three closed-source and three open-source models against 485 human participants, the study found MLLMs lack uniform SciVis literacy. Gemini emerged as the strongest model, surpassing the human mean, while open-source models performed below human baseline. Models excelled in scientific illustration, search, and spatial understanding tasks but struggled significantly with texture-based and integration-based visualizations, and quantitative estimation. Error analysis revealed consistent failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings underscore SciVis literacy as a critical, necessary benchmark dimension for multimodal AI system evaluation.

Key takeaway

For Machine Learning Engineers deploying MLLMs for scientific visualization interpretation, understand that current models exhibit highly uneven literacy. While Gemini shows promise, exceeding human performance, open-source alternatives generally fall short. You should carefully select models based on task type, utilizing MLLMs for spatial understanding or scientific illustration, but exercising caution and implementing human oversight for tasks requiring fine-grained quantitative estimation or interpreting complex texture-based visualizations.

Key insights

MLLMs show uneven scientific visualization literacy, with Gemini outperforming humans while open-source models lag.

Principles

Method

Benchmarked 6 MLLMs on a 49-item SciVis literacy test, comparing against 485 human participants under a closed-world protocol.

In practice

Topics

Code references

Best for: AI Engineer, AI Scientist, Machine Learning Engineer, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.