3 papers
cs.CV2025
A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
Alexey Magay, Dhurba Tripathi, Yu Hao +1
People with blindness and low vision (pBLV) face significant challenges, struggling to navigate environments and locate objects due to limited visual cues. Spatial reasoning is cru…
cs.CV2024
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
Yixiong Fang, Ziran Yang, Zhaorun Chen +2
Large vision-language models (LVLMs) excel at multimodal tasks but are prone to misinterpreting visual inputs, often resulting in hallucinations and unreliable outputs. We present…
cs.RO2024
Socially-Aware Robot Navigation Enhanced by Bidirectional Natural Language Conversations Using Large Language Models
Congcong Wen, Yifan Liu, Geeta Chandra Raju Bethala +7
Robot navigation is crucial across various domains, yet traditional methods focus on efficiency and obstacle avoidance, often overlooking human behavior in shared spaces. With the…