Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists
Summary
A study interviewed five blind or low-vision (BLV) and five sighted scientists from various STEM fields to understand their use of AI tools, specifically ChatGPT and Gemini, for querying multimodal scientific documents. The research characterized how scientists review multimodal content, including current practices and accessibility workarounds for visuals. It also gathered feedback on the suitability of AI-generated responses to multimodal queries. A key finding revealed that vague or incomplete image descriptions and incorrect AI outputs frequently cause both BLV and sighted scientists to abandon AI workflows. To support future research, the study contributes a dataset of 115 queries and responses from participant interactions with these AI tools on papers relevant to their fields. Implications for AI-powered scientific QA systems, emphasizing access across abilities and domains, are discussed.
Key takeaway
For AI Scientists developing multimodal scientific QA systems, you must prioritize robust accuracy and comprehensive accessibility features. Vague image descriptions and incorrect AI outputs significantly deter users, regardless of visual ability. Focus on improving AI model reliability for visual content interpretation and generating precise responses. Additionally, integrate user feedback from diverse populations, including blind and low-vision scientists, early in your development cycle to ensure equitable access and prevent workflow abandonment.
Key insights
AI-powered multimodal QA for scientific papers faces significant accessibility and accuracy challenges for both BLV and sighted scientists.
Principles
- Vague descriptions and incorrect AI outputs hinder workflow.
- Multimodal content access requires robust AI QA.
- Accessibility considerations are crucial for AI scientific tools.
Method
The study interviewed five BLV and five sighted scientists across STEM fields, observing their use of ChatGPT and Gemini to query multimodal scientific documents, and collected 115 queries/responses.
In practice
- Improve AI output accuracy for multimodal queries.
- Enhance image description quality for accessibility.
- Design AI QA systems for diverse user abilities.
Topics
- Multimodal AI
- Scientific QA Systems
- Accessibility
- Blind and Low-Vision
- ChatGPT
- Gemini
- User Studies
Best for: AI Product Manager, Research Scientist, AI Scientist, AI Ethicist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.