Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions
arXiv:2507.15692 · doi:10.1145/3663547.3746393
Abstract
Multimodal large language models (MLLMs) provide new opportunities for blind and low vision (BLV) people to access visual information in their daily lives. However, these models often produce errors that are difficult to detect without sight, posing safety and social risks in scenarios from medication identification to outfit selection. While BLV MLLM users use creative workarounds such as cross-checking between tools and consulting sighted individuals, these approaches are often time-consuming and impractical. We explore how systematically surfacing variations across multiple MLLM responses can support BLV users to detect unreliable information without visually inspecting the image. We contribute a design space for eliciting and presenting variations in MLLM descriptions, a prototype system implementing three variation presentation styles, and findings from a user study with 15 BLV participants. Our results demonstrate that presenting variations significantly increases users' ability to identify unreliable claims (by 4.9x using our approach compared to single descriptions) and significantly decreases perceived reliability of MLLM responses. 14 of 15 participants preferred seeing variations of MLLM responses over a single description, and all expressed interest in using our system for tasks from understanding a tornado's path to posting an image on social media.
18 pages, 6 figures
References in corpus (27)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- A Survey on Multimodal Large Language Models
- Design Principles for Generative AI Applications
- Language Models (Mostly) Know What They Know
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models
- Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
- Graphologue: Exploring Large Language Model Responses with Interactive Diagrams
- AI and Accessibility: A Discussion of Ethical Considerations
- Writer-Defined AI Personas for On-Demand Feedback Generation
- WorldScribe: Towards Context-Aware Live Visual Descriptions
- AVscript: Accessible Video Editing with Audio-Visual Scripts
- Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People
- Making Short-Form Videos Accessible with Hierarchical Video Summaries
- One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
- RELIC: Investigating Large Language Model Responses using Self-Consistency
- Misfitting With AI: How Blind People Verify and Contest AI Errors
- ImageAssist: Tools for Enhancing Touchscreen-Based Image Exploration Systems for Blind and Low Vision Users
- Explanations Can Reduce Overreliance on AI Systems During Decision-Making
- Context-Aware Image Descriptions for Web Accessibility
- TADA: Making Node-link Diagrams Accessible to Blind and Low-Vision People
- WanderGuide: Indoor Map-less Robotic Guide for Exploration by Blind People
- Understanding How Blind Users Handle Object Recognition Errors: Strategies and Challenges
- EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
- VisiMark: Characterizing and Augmenting Landmarks for People with Low Vision in Augmented Reality to Support Indoor Navigation
- Personalized Language Modeling from Personalized Human Feedback
- Uncertainty quantification by direct propagation of shallow ensembles