collaborators

7 papers

cs.CV2026

Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation

Muyi Bao, Yuxin Cai, Hang Xu +7

Vision-language models (VLMs) have become a common foundation for vision-and-language navigation in continuous environments (VLN-CE). Yet most VLM-based methods cast navigation as…

cs.RO2026

IntentNav: Learning Spatial-Visual Object Navigation from Human Demonstrations

Yuxin Cai, Zongtai Li, Maonan Wang +9

Object navigation requires a robot to search for an unobserved target in an unknown environment by deciding where to explore next under partial observability. Effective search rese…

cs.RO2026

SysNav: Multi-Level Systematic Cooperation Enables Real-World, Cross-Embodiment Object Navigation

Haokun Zhu, Zongtai Li, Zihan Liu +8

Object navigation (ObjectNav) in real-world environments is a complex problem that requires simultaneously addressing multiple challenges, including complex spatial structure, long…

cs.CV2025

LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models

Pranav Saxena, Avigyan Bhattacharya, Ji Zhang +1

Referential grounding in outdoor driving scenes is challenging due to large scene variability, many visually similar objects, and dynamic elements that complicate resolving natural…

cs.RO2025

STRIVE: Structured Representation Integrating VLM Reasoning for Efficient Object Navigation

Haokun Zhu, Zongtai Li, Zhixuan Liu +4

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs t…

cs.CV2025

SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models

Nader Zantout, Haochen Zhang, Pujith Kachana +4

Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is…