6 papers · 1 filter
TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation
Jianyi Zhou, Feiyang Hong, Yunhao Li +9
Dexterous manipulation in everyday environments requires both anticipation and reaction: a robot must predict how contact should evolve while rapidly correcting local errors caused…
MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation
Jia Zheng, Teli Ma, Yudong Fan +3
World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional vi…
TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video
Jianyi Zhou, Ziteng Gao, Feiyang Hong +11
Egocentric human video data, which captures rich human-environment interactions and can be collected at scale, has become a key driver of embodied intelligence research. However, e…
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Teli Ma, Jia Zheng, Zifan Wang +4
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretrainin…
Hybrid Diffusion Policies with Projective Geometric Algebra for Efficient Robot Manipulation Learning
Xiatao Sun, Yuxuan Wang, Shuo Yang +2
Diffusion policies are a powerful paradigm for robot learning, but their training is often inefficient. A key reason is that networks must relearn fundamental spatial concepts, suc…
Dynamic Rank Adjustment in Diffusion Policies for Efficient and Flexible Training
Xiatao Sun, Shuo Yang, Yinxing Chen +3
Diffusion policies trained via offline behavioral cloning have recently gained traction in robotic motion generation. While effective, these policies typically require a large numb…