collaborators

9 papers

cs.RO2026

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Tao Lin, Yuxin Du, Yiran Mao +13

Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…

cs.CV2026

Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference

Ziyan Liu, Yeqiu Chen, Hongyi Cai +4

Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and action execution. In real-time deployme…

cs.RO2026

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Runze Wang, Yuqian Fu, Yu Li +7

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determ…

cs.CV2026

Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model

Tao Lin, Yuxin Du, Jiting Liu +14

Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often s…

cs.CV2026

Focusable Monocular Depth Estimation

Yuxin Du, Tao Lin, Zile Zhong +7

Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-…

cs.RO2026

RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction

Zewei Ye, Weifeng Lu, Minghao Ye +4

Vision-Language-Action (VLA) models have recently advanced robotic manipulation by translating natural-language instructions and visual observations into control actions. However,…