DRAGON: A Dialogue-Based Robot for Assistive Navigation with Visual Language Grounding
arXiv:2307.06924 · doi:10.1109/LRA.2024.3362591
Abstract
Persons with visual impairments (PwVI) have difficulties understanding and navigating spaces around them. Current wayfinding technologies either focus solely on navigation or provide limited communication about the environment. Motivated by recent advances in visual-language grounding and semantic navigation, we propose DRAGON, a guiding robot powered by a dialogue system and the ability to associate the environment with natural language. By understanding the commands from the user, DRAGON is able to guide the user to the desired landmarks on the map, describe the environment, and answer questions from visual observations. Through effective utilization of dialogue, the robot can ground the user's free-form descriptions to landmarks in the environment, and give the user semantic information through spoken language. We conduct a user study with blindfolded participants in an everyday indoor environment. Our results demonstrate that DRAGON is able to communicate with the user smoothly, provide a good guiding experience, and connect users with their surrounding environment in an intuitive manner. Videos and code are available at https://sites.google.com/view/dragon-wayfinding/home.
Published in IEEE Robotics and Automation Letters (RA-L)
References in corpus (5)
- Robust Speech Recognition via Large-Scale Weak Supervision
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- DIET: Lightweight Language Understanding for Dialogue Systems
- LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
- "I am the follower, also the boss": Exploring Different Levels of Autonomy and Machine Forms of Guiding Robots for the Visually Impaired