most citedUni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

2 citations · 2 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2025

4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer

Xianfeng Wu, Yajing Bai, Minghan Li +5

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic envi…

cs.RO2025

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

Anqi Li, Zhiyong Wang, Jiazhao Zhang +5

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This ta…

cs.RO2025

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

Jiahang Liu, Yunpeng Qi, Jiazhao Zhang +9

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously…

cs.RO2025

Embodied Navigation Foundation Model

Jiazhao Zhang, Anqi Li, Yunpeng Qi +14

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions.…

cs.RO2025

TrackVLA: Embodied Visual Tracking in the Wild

Shaoan Wang, Jiazhao Zhang, Minghan Li +7

Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inh…

cs.RO20242 cited

Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

Jiazhao Zhang, Kunyu Wang, Shaoan Wang +6

A practical navigation agent must be capable of handling a wide range of interaction demands, such as following instructions, searching objects, answering questions, tracking peopl…