2 citations · 2 across the 4 of their papers we have counts for
4 papers
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
Jiahang Liu, Yunpeng Qi, Jiazhao Zhang +9
Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously…
Embodied Navigation Foundation Model
Jiazhao Zhang, Anqi Li, Yunpeng Qi +14
Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions.…
TrackVLA: Embodied Visual Tracking in the Wild
Shaoan Wang, Jiazhao Zhang, Minghan Li +7
Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inh…
Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks
Jiazhao Zhang, Kunyu Wang, Shaoan Wang +6
A practical navigation agent must be capable of handling a wide range of interaction demands, such as following instructions, searching objects, answering questions, tracking peopl…