collaborators

10 papers

cs.CV2026

Learning Additively Compositional Latent Actions for Embodied AI

Hangxing Wei, Xiaoyu Chen, Chuheng Zhang +5

Latent action learning infers pseudo-action labels from visual transitions, providing an approach to leverage internet-scale video for embodied AI. However, most methods learn late…

cs.RO2025

Discover, Learn, and Reinforce: Scaling Vision-Language-Action Pretraining with Diverse RL-Generated Trajectories

Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang +6

Scaling vision-language-action (VLA) model pre-training requires large volumes of diverse, high-quality manipulation trajectories. Most current data is obtained via human teleopera…

cs.RO2025

How Do VLAs Effectively Inherit from VLMs?

Chuheng Zhang, Rushuai Yang, Xiaoyu Chen +4

Vision-language-action (VLA) models hold the promise to attain generalizable embodied control. To achieve this, a pervasive paradigm is to leverage the rich vision-semantic priors…

cs.LG2025

Co-Evolving Latent Action World Models

Yucen Wang, Fengming Zhang, De-Chuan Zhan +3

Adapting pretrained video generation models into controllable world models via latent actions is a promising step towards creating generalist world models. The dominant paradigm ad…

cs.RO2025

Beyond Human Demonstrations: Diffusion-Based Reinforcement Learning to Generate Data for VLA Training

Rushuai Yang, Hangxing Wei, Ran Zhang +8

Vision-language-action (VLA) models have shown strong generalization across tasks and embodiments; however, their reliance on large-scale human demonstrations limits their scalabil…

cs.CV2025

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models

Jiansong Wan, Chengming Zhou, Jinkua Liu +14

Recent studies have explored pretrained (foundation) models for vision-based robotic navigation, aiming to achieve generalizable navigation and positive transfer across diverse env…