7 papers
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Yen-Jen Wang, Jiaman Li, Sirui Chen +9
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learning this mapping requires synchronized egoc…
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
Sirui Chen, Shibo Zhao, Zhen Wu +3
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracking cannot resolve the physical ambiguitie…
AnyLift: Scaling Motion Reconstruction from Internet Videos via 2D Diffusion
Hongjie Li, Heng Yu, Jiaman Li +4
Reconstructing 3D human motion and human-object interactions (HOI) from Internet videos is a fundamental step toward building large-scale datasets of human behavior. Existing metho…
WHOLE: World-Grounded Hand-Object Lifted from Egocentric Videos
Yufei Ye, Jiaman Li, Ryan Rong +1
Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. C…
Human-Object Interaction from Human-Level Instructions
Zhen Wu, Jiaman Li, Pei Xu +1
Intelligent agents must autonomously interact with the environments to perform daily tasks based on human-level instructions. They need a foundational understanding of the world to…
Lifting Motion to the 3D World via 2D Diffusion
Jiaman Li, C. Karen Liu, Jiajun Wu
Estimating 3D motion from 2D observations is a long-standing research challenge. Prior work typically requires training on datasets containing ground truth 3D motions, limiting the…