1 citations · 1 across the 4 of their papers we have counts for
6 papers
EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments
Xilin Wang, Guoxi Zhang, Hongming Xu +3
Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exp…
UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation
Zhuofan Zhang, Tianxu Wang, Guoxi Zhang +6
Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guar…
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation
Ziyu Zhu, Xilin Wang, Yixuan Li +9
Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world.…
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
Tianxu Wang, Zhuofan Zhang, Ziyu Zhu +5
3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplore…
Task-oriented Sequential Grounding and Navigation in 3D Scenes
Zhuofan Zhang, Ziyu Zhu, Junhao Li +8
Grounding natural language in 3D environments is a critical step toward achieving robust 3D vision-language alignment. Current datasets and models for 3D visual grounding predomina…
Unifying 3D Vision-Language Understanding via Promptable Queries
Ziyu Zhu, Zhuofan Zhang, Xiaojian Ma +6
A unified model for 3D vision-language (3D-VL) understanding is expected to take various scene representations and perform a wide range of tasks in a 3D scene. However, a considera…