11 papers
Bespoke Visual Assistance: What and How do Blind and Low-Vision People Create with Agentic Programming?
Ellie Seehorn, Gene S-H Kim, Aziz Zeidieh +6
AI-powered assistive technologies have long supported blind and low vision (BLV) people in everyday tasks, but they are general-purpose and often fall short of meeting complex, ind…
Sidekick: Designing Communication for Effective Multitasking with Computer Use Agents
Ruei-Che Chang, Wenqian Xu, Dingzeyu Li +2
Computer Use Agents (CUAs) can autonomously execute complex, multi-step tasks within GUIs, enhancing efficiency through parallel multitasking. However, our formative studies with C…
Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
Yayuan Li, Chenglin Li, Jingying Wang +3
Instructional videos are the dominant medium for learning physical tasks, yet they rarely match the user's real-world visual context. Motor simulation and cognitive load theories p…
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
Chen Liang, Xirui Jiang, Naihao Deng +2
AI agents operating on user interfaces must understand how interfaces communicate state and feedback to act reliably. As a core communicative modality, animations are increasingly…
StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
Ruei-Che Chang, Xirui Jiang, Rosiana Natalie +6
Real-world environments evolve continuously, yet blind and low-vision (BLV) individuals often have limited access to understanding how they change over time. Unexpected or relocate…
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
Ananya Gubbi Mohanbabu, Rosiana Natalie, Brandon Kim +2
Computer Use Agents (CUAs) operate interfaces by pointing, clicking, and typing -- mirroring interactions of sighted users (SUs) who can thus monitor CUAs and share control. CUAs d…