9 papers
LA4VLA: Learning to Act without Seeing via Language-Action Pretraining
Tao Lin, Yuxin Du, Yiran Mao +13
Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…
Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
Ziyan Liu, Yeqiu Chen, Hongyi Cai +4
Vision-Language-Action (VLA) models have shown great potential for embodied AI by integrating visual perception, language understanding, and action execution. In real-time deployme…
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
Runze Wang, Yuqian Fu, Yu Li +7
Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determ…
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Tao Lin, Yuxin Du, Jiting Liu +14
Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often s…
Focusable Monocular Depth Estimation
Yuxin Du, Tao Lin, Zile Zhong +7
Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-…
RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction
Zewei Ye, Weifeng Lu, Minghao Ye +4
Vision-Language-Action (VLA) models have recently advanced robotic manipulation by translating natural-language instructions and visual observations into control actions. However,…