activity
20242026
most citedTask-oriented Sequential Grounding and Navigation in 3D Scenes

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.RO2026

EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments

Xilin Wang, Guoxi Zhang, Hongming Xu +3

Lifelong navigation (LN) requires an embodied agent to solve a sequence of navigation subtasks in the same environment. Since solving each subtask from scratch incurs redundant exp…

cs.RO2026

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

Zhuofan Zhang, Tianxu Wang, Guoxi Zhang +6

Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity of the target does not guar…

cs.CV2025

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

Ziyu Zhu, Xilin Wang, Yixuan Li +9

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world.…

cs.CV2025

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu +5

3D visual grounding has made notable progress in localizing objects within complex 3D scenes. However, grounding referring expressions beyond objects in 3D scenes remains unexplore…

cs.CV2024★ 1 cited

Task-oriented Sequential Grounding and Navigation in 3D Scenes

Zhuofan Zhang, Ziyu Zhu, Junhao Li +8

Grounding natural language in 3D environments is a critical step toward achieving robust 3D vision-language alignment. Current datasets and models for 3D visual grounding predomina…

cs.CV2024

Unifying 3D Vision-Language Understanding via Promptable Queries

Ziyu Zhu, Zhuofan Zhang, Xiaojian Ma +6

A unified model for 3D vision-language (3D-VL) understanding is expected to take various scene representations and perform a wide range of tasks in a 3D scene. However, a considera…