From the 1 of 10 linked papers with an AI index.
10 papers
Motubrain: An Advanced World Action Model for Robot Control
Motubrain Team, Chendong Xiang, Fan Bao +17
Motubrain is a unified world action model that jointly learns video and robot actions using a UniDiffuser and Mixture-of-Transformers architecture, enabling policy learning, world…
AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation
Hengkai Tan, Yao Feng, Xinyi Mao +5
Learning generalizable manipulation policies hinges on data, yet robot manipulation data is scarce and often entangled with specific embodiments, making both cross-task and cross-p…
Motus: A Unified Latent Action World Model
Hongzhe Bi, Hengkai Tan, Shenghao Xie +13
While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation pr…
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
Yao Feng, Hengkai Tan, Xinyi Mao +5
Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action p…
Vidarc: Embodied Video Diffusion Model for Closed-loop Control
Yao Feng, Chendong Xiang, Xinyi Mao +7
Robotic arm manipulation in data-scarce settings is a highly challenging task due to the complex embodiment dynamics and diverse contexts. Recent video-based approaches have shown…
GentleHumanoid: Learning Upper-body Compliance for Contact-rich Human and Object Interaction
Qingzhou Lu, Yao Feng, Baiyu Shi +3
Humanoid robots are expected to operate in human-centered environments where safe and natural physical interaction is essential. However, most recent reinforcement learning (RL) po…