paper

Sighted by Default: Addressing Implicit Vision Assumptions in Real-Time VLM Assistance for BLV Users

arXiv:2511.00945 · doi:10.1145/3830398.3830496

Abstract

Vision-Language Model (VLM)-based assistance is reshaping independence for blind and low-vision (BLV) users, yet current tools fail in dynamic settings. While request-response architectures like BeMyAI impose prohibitive latencies, our formative study (N = 15) reveals that even real-time alternatives like Doubao fail due to a deeper structural problem: sighted-default bias -- the implicit assumption that users possess parallel visual access to their surroundings. This bias manifests as verbose, vision-centric narratives that overlook the serial nature of auditory perception, flooding the user's limited cognitive bandwidth with information that is neither timely nor actionable. To address this, we derive three design principles: Continuity, Conciseness, and Calibrated Honesty. We present VIA-Agent, which co-optimizes a specialized cognitive core with a low-latency Real-Time Communication (RTC) architecture for continuous bidirectional streaming. In a within-subjects evaluation (N = 9), VIA-Agent matched Doubao's success rate, significantly reducing mean task time by 21.4% (91.7s vs. 116.7s) and conversational turns from 5.9 to 4.3 while achieving higher trust.

Accepted to UIST 2026