4 papers
Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models
Yuhang Song, Bor-Jiun Lin, Jiaxu Liu +3
Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independen…
Optimizing Neurorobot Policy under Limited Demonstration Data through Preference Regret
Viet Dung Nguyen, Yuhang Song, Anh Nguyen +3
Robot reinforcement learning from demonstrations (RLfD) assumes that expert data is abundant; this is usually unrealistic in the real world given data scarcity as well as high coll…
GuideTWSI: A Diverse Tactile Walking Surface Indicator Dataset from Synthetic and Real-World Images for Blind and Low-Vision Navigation
Hochul Hwang, Soowan Yang, Anh N. H. Nguyen +6
Tactile Walking Surface Indicators (TWSIs) are safety-critical landmarks that blind and low-vision (BLV) pedestrians use to locate crossings and hazard zones. From our observation…
Fine-Grained Alignment in Vision-and-Language Navigation through Bayesian Optimization
Yuhang Song, Mario Gianni, Chenguang Yang +4
This paper addresses the challenge of fine-grained alignment in Vision-and-Language Navigation (VLN) tasks, where robots navigate realistic 3D environments based on natural languag…