InterGen: Diffusion-based Multi-human Motion Generation under Complex Interactions
arXiv:2304.05684 · doi:10.1007/s11263-024-02042-6
Abstract
We have recently seen tremendous progress in diffusion advances for generating realistic human motions. Yet, they largely disregard the multi-human interactions. In this paper, we present InterGen, an effective diffusion-based approach that incorporates human-to-human interactions into the motion diffusion process, which enables layman users to customize high-quality two-person interaction motions, with only text guidance. We first contribute a multimodal dataset, named InterHuman. It consists of about 107M frames for diverse two-person interactions, with accurate skeletal motions and 23,337 natural language descriptions. For the algorithm side, we carefully tailor the motion diffusion model to our two-person interaction setting. To handle the symmetry of human identities during interactions, we propose two cooperative transformer-based denoisers that explicitly share weights, with a mutual attention mechanism to further connect the two denoising processes. Then, we propose a novel representation for motion input in our interaction diffusion model, which explicitly formulates the global relations between the two performers in the world frame. We further introduce two novel regularization terms to encode spatial relations, equipped with a corresponding damping scheme during the training of our interaction diffusion model. Extensive experiments validate the effectiveness and generalizability of InterGen. Notably, it can generate more diverse and compelling two-person motions than previous methods and enables various downstream applications for human interactions.
accepted by IJCV 2024
References in corpus (10)
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- NTU RGB+D 120: A Large-Scale Benchmark for 3D Human Activity Understanding
- Score-Based Generative Modeling through Stochastic Differential Equations
- Improved Denoising Diffusion Probabilistic Models
- Action2Motion: Conditioned Generation of 3D Human Motions
- Robust Motion In-betweening
- STAR: Sparse Trained Articulated Human Body Regressor
- Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
- LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR Sensors
- HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes
Cited by in corpus (4)
- A Survey on Human Interaction Motion Generation
- InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios
- BOTH2Hands: Inferring 3D Hands from Both Text Prompts and Body Dynamics
- InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation