activity
20242026
most citedChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

2 citations · 2 across the 3 of their papers we have counts for

collaborators
Showing cs.ROShow all

5 papers · 1 filter

cs.RO2026

VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models

Zirui Ge, Pengxiang Ding, Baohua Yin +16

Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…

cs.RO20252 cited

ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model

Zhongyi Zhou, Yichen Zhu, Minjie Zhu +8

Humans possess a unified cognitive ability to perceive, comprehend, and interact with the physical world. Why can't large language models replicate this holistic understanding? Thr…

cs.RO2024

CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance

Jinming Li, Yichen Zhu, Zhibin Tang +8

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving…

cs.RO2024

Scaling Diffusion Policy in Transformer to 1 Billion Parameters for Robotic Manipulation

Minjie Zhu, Yichen Zhu, Jinming Li +8

Diffusion Policy is a powerful technique tool for learning end-to-end visuomotor robot control. It is expected that Diffusion Policy possesses scalability, a key attribute for deep…

cs.RO2024

TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

Junjie Wen, Yichen Zhu, Jinming Li +9

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA…