3 papers
cs.AI2026
Look on Demand: A Cognitive Scheduling Framework for Visual Evidence Acquisition in Multimodal Reasoning
Yang Zhang, Xiaoshuai Sun, Rui Zhao +5
Existing multimodal reasoning approaches predominantly follow two paradigms: converting visual inputs into text prior to reasoning, or performing end-to-end reasoning within a unif…
cs.HC2025
"Does the cafe entrance look accessible? Where is the door?" Towards Geospatial AI Agents for Visual Inquiries
Jon E. Froehlich, Jared Hwang, Zeyu Wang +7
Interactive digital maps have revolutionized how people travel and learn about the world; however, they rely on pre-existing structured data in GIS databases (e.g., road networks,…
cs.HC2025
Accessibility Scout: Personalized Accessibility Scans of Built Environments
William Huang, Xia Su, Jon E. Froehlich +1
Assessing the accessibility of unfamiliar built environments is critical for people with disabilities. However, manual assessments, performed by users or their personal health prof…