Transflower: probabilistic autoregressive dance generation with multimodal attention
arXiv:2106.13871 · doi:10.1145/3478513.3480570
Abstract
Dance requires skillful composition of complex movements that follow rhythmic, tonal and timbral features of music. Formally, generating dance conditioned on a piece of music can be expressed as a problem of modelling a high-dimensional continuous motion signal, conditioned on an audio signal. In this work we make two contributions to tackle this problem. First, we present a novel probabilistic autoregressive architecture that models the distribution over future poses with a normalizing flow conditioned on previous poses as well as music context, using a multimodal transformer encoder. Second, we introduce the currently largest 3D dance-motion dataset, obtained with a variety of motion-capture technologies, and including both professional and casual dancers. Using this dataset, we compare our new model against two baselines, via objective metrics and a user study, and show that both the ability to model a probability distribution, as well as being able to attend over a large motion and music context are necessary to produce interesting, diverse, and realistic dance that matches the music.
Article presented at SIGGRAPH Asia 2021, and published in ACM Transactions on Graphics
References in corpus (12)
- Language Models are Few-Shot Learners
- Scaling Laws for Neural Language Models
- Zero-Shot Text-to-Image Generation
- AMP: Adversarial Motion Priors for Stylized Physics-Based Character Control
- Character Controllers Using Motion VAEs
- Flow++: Improving Flow-Based Generative Models with Variational Dequantization and Architecture Design
- Scaling Laws for Autoregressive Generative Modeling
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Learning to Generate Diverse Dance Motions with Transformer
- A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020
- ChoreoNet: Towards Music to Dance Synthesis with Choreographic Action Unit
- Dancing to Music
Cited by in corpus (12)
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
- Transformer Inertial Poser: Real-time Human Motion Reconstruction from Sparse IMUs with Simultaneous Terrain Generation
- The Ethical Implications of Generative Audio Models: A Systematic Literature Review
- DanceGen: Supporting Choreography Ideation and Prototyping with Generative AI
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
- Evaluation of Generative Models for Emotional 3D Animation Generation in VR
- A Two-part Transformer Network for Controllable Motion Synthesis
- DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis
- Model See Model Do: Speech-Driven Facial Animation with Style Control
- STyMo: Fast and Controllable Few-Shot Motion Style Transfer