collaborators

6 papers

cs.RO2026

Learning Transferable Dynamics Priors from Action to World Modeling

Ze Huang, Jiahui Zhang, Hairuo Liu +3

We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual sc…

cs.CV2026

Dual-Domain Representation Alignment: Bridging 2D and 3D Vision via Geometry-Aware Architecture Search

Haoyu Zhang, Zhihao Yu, Rui Wang +3

Modern computer vision requires balancing predictive accuracy with real-time efficiency, yet the high inference cost of large vision models (LVMs) limits deployment on resource-con…

cs.RO2026

VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models

Zirui Ge, Pengxiang Ding, Baohua Yin +16

Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…

cs.RO2025

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Jinming Li, Yichen Zhu, Zhibin Tang +8

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving…

cs.RO2025

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Junjie Wen, Yichen Zhu, Jinming Li +9

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA…

cs.RO2025

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Zhongyi Zhou, Yichen Zhu, Minjie Zhu +8

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Thr…