2 citations · 3 across the 2 of their papers we have counts for
5 papers · 1 filter
Chain-of-Modality: Learning Manipulation Programs from Multimodal Human Videos with Vision-Language-Models
Chen Wang, Fei Xia, Wenhao Yu +6
Learning to perform manipulation tasks from human videos is a promising approach for teaching robots. However, many manipulation tasks require changing control parameters during ta…
IPPON: Common Sense Guided Informative Path Planning for Object Goal Navigation
Kaixian Qu, Jie Tan, Tingnan Zhang +3
Navigating efficiently to an object in an unexplored environment is a critical skill for general-purpose intelligent robots. Recent approaches to this object goal navigation proble…
Mobility VLA: Multimodal Instruction Navigation with Long-Context VLMs and Topological Graphs
Hao-Tien Lewis Chiang, Zhuo Xu, Zipeng Fu +19
An elusive goal in navigation research is to build an intelligent agent that can understand multimodal instructions including natural language and image, and perform useful navigat…
CoNVOI: Context-aware Navigation using Vision Language Models in Outdoor and Indoor Environments
Adarsh Jagan Sathyamoorthy, Kasun Weerakoon, Mohamed Elnoor +6
We present ConVOI, a novel method for autonomous robot navigation in real-world indoor and outdoor environments using Vision Language Models (VLMs). We employ VLMs in two ways: fir…
Creative Robot Tool Use with Large Language Models
Mengdi Xu, Peide Huang, Wenhao Yu +7
Tool use is a hallmark of advanced intelligence, exemplified in both animal behavior and robotic capabilities. This paper investigates the feasibility of imbuing robots with the ab…