12 papers
Lift3D-VLA: Lifting VLA Models to 3D Geometry and Dynamics-Aware Manipulation
Jiaming Liu, Qingpo Wuwu, Nuowei Han +8
Recently, Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse tasks. However, effective robotic manipulation in physical environments fundame…
LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation
Jiaming Liu, Yinxi Wang, Chenyang Gu +15
Human-hand demonstrations provide a direct and scalable source of physical interaction data for robot learning. While manual retargeting is indispensable for establishing kinematic…
MV-WAM: Manifold-Aware World Action Model with Value Augmentation
Jintao Chen, Peidong Jia, Qingpo Wuwu +13
Achieving robust and generalizable manipulation across diverse environments remains a fundamental challenge in embodied robotics. Recent world action models achieve strong in-domai…
LaST: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model
Zhuoyang Liu, Jiaming Liu, Hao Chen +11
Vision-Language-Action (VLA) models have recently shown strong generalization, with some approaches seeking to explicitly generate linguistic reasoning traces or predict future obs…
HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
Qiuxuan Feng, Jiale Yu, Jiaming Liu +8
World Action Models (WAMs) have emerged as a promising paradigm for robot control by modeling physical dynamics. Current WAMs generally follow two paradigms: the "Imagine-then-Exec…
CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference
Zhitong Dong, Chao Li, Jie Yu +1
Aesthetic image cropping aims to enhance the aesthetic quality of an image by improving its composition through spatial cropping. Previous methods often rely on saliency prediction…