most citedUni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

2 citations · 2 across the 5 of their papers we have counts for

collaborators

11 papers

cs.RO2025

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

Anqi Li, Zhiyong Wang, Jiazhao Zhang +5

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This ta…

cs.RO2025

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

Jiahang Liu, Yunpeng Qi, Jiazhao Zhang +9

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously…

cs.RO2025

Embodied Navigation Foundation Model

Jiazhao Zhang, Anqi Li, Yunpeng Qi +14

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions.…

cs.CV2025

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Wenyao Zhang, Hongsi Liu, Zekun Qi +11

Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot ma…

cs.CV2025

DexVLG: Dexterous Vision-Language-Grasp Model at Scale

Jiawei He, Danshi Li, Xinqiang Yu +7

As large models gain traction, vision-language-action (VLA) systems are enabling robots to tackle increasingly complex tasks. However, limited by the difficulty of data collection,…

cs.RO2025

TrackVLA: Embodied Visual Tracking in the Wild

Shaoan Wang, Jiazhao Zhang, Minghan Li +7

Embodied visual tracking is a fundamental skill in Embodied AI, enabling an agent to follow a specific target in dynamic environments using only egocentric vision. This task is inh…