Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
arXiv:2211.09707 · doi:10.1145/3592458
Abstract
Diffusion models have experienced a surge of interest as highly expressive yet efficiently trainable probabilistic models. We show that these models are an excellent fit for synthesising human motion that co-occurs with audio, e.g., dancing and co-speech gesticulation, since motion is complex and highly ambiguous given audio, calling for a probabilistic description. Specifically, we adapt the DiffWave architecture to model 3D pose sequences, putting Conformers in place of dilated convolutions for improved modelling power. We also demonstrate control over motion style, using classifier-free guidance to adjust the strength of the stylistic expression. Experiments on gesture and dance generation confirm that the proposed method achieves top-of-the-line motion quality, with distinctive styles whose expression can be made more or less pronounced. We also synthesise path-driven locomotion using the same model architecture. Finally, we generalise the guidance procedure to obtain product-of-expert ensembles of diffusion models and demonstrate how these may be used for, e.g., style interpolation, a contribution we believe is of independent interest. See https://www.speech.kth.se/research/listen-denoise-action/ for video examples, data, and code.
20 pages, 9 figures. Published in ACM ToG and presented at SIGGRAPH 2023
References in corpus (10)
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- Character Controllers Using Motion VAEs
- Analyzing Input and Output Representations for Speech-Driven Gesture Generation
- Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
- The GENEA Challenge 2022: A large evaluation of data-driven co-speech gesture generation
- A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020
- ChoreoNet: Towards Music to Dance Synthesis with Choreographic Action Unit
- Generating coherent spontaneous speech and gesture from text
- Integrated Speech and Gesture Synthesis
- Robust Classification using Hidden Markov Models and Mixtures of Normalizing Flows
Cited by in corpus (19)
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- DanceGen: Supporting Choreography Ideation and Prototyping with Generative AI
- Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
- Diffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation
- GaussianPrediction: Dynamic 3D Gaussian Prediction for Motion Extrapolation and Free View Synthesis
- ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAE
- LGTM: Local-to-Global Text-Driven Human Motion Diffusion Model
- Exploring AI-assisted Ideation and Prototyping for Choreography
- Evaluation of Generative Models for Emotional 3D Animation Generation in VR
- Decoupling Contact for Fine-Grained Motion Style Transfer
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
- MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
- Fatigue-PINN: Physics-Informed Fatigue-Driven Motion Modulation and Synthesis
- A Survey on Human Interaction Motion Generation
- ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition
- MAMM: Motion Control via Metric-Aligning Motion Matching
- Real-time Diverse Motion In-betweening with Space-time Control
- Multi-Resolution Generative Modeling of Human Motion from Limited Data
- InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios