activity
20242026
collaborators
Showing cs.ROShow all

18 papers · 1 filter

cs.RO2026

Data Pyramid for Embodied Manipulation: A Survey

Yifan Ye, Yankai Fu, Yaoxu Lv +26

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations w…

cs.RO2026

Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation

Jiaming Liu, Qingpo Wuwu, Nuowei Han +8

Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundame…

cs.RO2026

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

Shengbang Liu, Yueru Jia, Yuyang Yan +7

Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still struggle with contact-rich tasks, where minor contact perturbations…

cs.RO2026

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation

Jiaming Liu, Yinxi Wang, Chenyang Gu +15

Human-hand demonstrations provide a direct and scalable source of physical interaction data for robot learning. While manual retargeting is indispensable for establishing kinematic…

cs.RO2026

MV-WAM: Manifold-Aware World Action Model with Value Augmentation

Jintao Chen, Peidong Jia, Qingpo Wuwu +13

Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domai…

cs.RO2026

LaST: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model

Zhuoyang Liu, Jiaming Liu, Hao Chen +11

Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future obs…