2 citations · 2 across the 5 of their papers we have counts for
8 papers · 1 filter
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
Anqi Li, Zhiyong Wang, Jiazhao Zhang +5
Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This ta…
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
Jiahang Liu, Yunpeng Qi, Jiazhao Zhang +9
Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously…
Embodied Navigation Foundation Model
Jiazhao Zhang, Anqi Li, Yunpeng Qi +14
Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions.…
TrackVLA: Embodied Visual Tracking in the Wild
Shaoan Wang, Jiazhao Zhang, Minghan Li +7
Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inh…
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Shengliang Deng, Mi Yan, Songlin Wei +10
Embodied foundation models are gaining increasing attention for their zero-shot generalization, scalability, and adaptability to new tasks through few-shot post-training. However,…
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
Zekun Qi, Wenyao Zhang, Yufei Ding +15
While spatial reasoning has made progress in object localization relationships, it often overlooks object orientation-a key factor in 6-DoF fine-grained manipulation. Traditional p…