9 papers
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Yifei Zhu, Mingyi Shi, Yangyang Cai +3
Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into…
Prior-First, Condition-Second: Scalable and Controllable Hand Motion Completion
Mingyi Shi, Xuelin Chen, Taku Komura
Synthesizing hand motion that matches the full body motion and the semantic labels is a difficult task due to their high degrees of freedom and the lack of semantic labels. To cope…
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
Wenjia Wang, Liang Pan, Huaijin Pi +8
Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting.…
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors
Xiaoyan Cong, Zekun Li, Zhiyang Dou +9
Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets…
CHOICE: Coordinated Human-Object Interaction in Cluttered Environments for Pick-and-Place Actions
Jintao Lu, He Zhang, Yuting Ye +3
Animating human-scene interactions such as pick-and-place tasks in cluttered, complex layouts is a challenging task, with objects of a wide variation of geometries and articulation…
InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios
Leo Ho, Yinghao Huang, Dafei Qin +5
We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on co…