collaborators

7 papers

cs.CV2026

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

Ziyan Liu, Yeqiu Chen, Hongyi Cai +4

Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and action execution. In real-time deployme…

cs.RO2026

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Runze Wang, Yuqian Fu, Yu Li +7

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determ…

cs.CV2026

R3D: Revisiting 3D Policy Learning

Zhengdong Hong, Shenrui Wu, Haozhe Cui +8

3D policy learning promises superior generalization and cross-embodiment transfer, but progress has been hindered by training instabilities and severe overfitting, precluding the a…

cs.RO2026

U-ARM : Ultra low-cost general teleoperation interface for robot manipulation

Yanwen Zou, Zhaoye Zhou, Chenyang Shi +4

We propose U-Arm, a low-cost and rapidly adaptable leader-follower teleoperation framework designed to interface with most of commercially available robotic arms. Our system suppor…

cs.RO2025

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

Tao Lin, Yilei Zhong, Yuxin Du +11

Vision-Language-Action (VLA) models have emerged as a powerful framework that unifies perception, language, and control, enabling robots to perform diverse tasks through multimodal…

cs.RO2025

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

Tao Lin, Gen Li, Yilei Zhong +5

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These model…