8 papers
TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation
Jianyi Zhou, Feiyang Hong, Yunhao Li +9
Dexterous manipulation in everyday environments requires both anticipation and reaction: a robot must predict how contact should evolve while rapidly correcting local errors caused…
MotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation
Jia Zheng, Teli Ma, Yudong Fan +3
World Action Models (WAMs) couple a video dynamics prior to the policy and have shown encouraging results on tabletop manipulation, but iterative denoising over high-dimensional vi…
TouchAnything: A Dataset and Framework for Bimanual Tactile Estimation from Egocentric Video
Jianyi Zhou, Ziteng Gao, Feiyang Hong +11
Egocentric human video data, which captures rich human-environment interactions and can be collected at scale, has become a key driver of embodied intelligence research. However, e…
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
Zheyu Zhang, Shuo Yang, Bardh Prenkaj +1
Generative tabular augmentation is appealing in data-scarce domains, yet the prevailing focus on distributional fidelity does not reliably translate into better downstream models.…
From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation
Bohan Li, Shuojue Yang, Baorui Peng +10
Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dimensional control vectors must prec…
DiT4DiT: Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control
Teli Ma, Jia Zheng, Zifan Wang +4
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretrainin…