Audio to Body Dynamics
arXiv:1712.09382 · doi:10.1109/CVPR.2018.00790
Abstract
We present a method that gets as input an audio of violin or piano playing, and outputs a video of skeleton predictions which are further used to animate an avatar. The key idea is to create an animation of an avatar that moves their hands similarly to how a pianist or violinist would do, just from audio. Aiming for a fully detailed correct arms and fingers motion is a goal, however, it's not clear if body movement can be predicted from music at all. In this paper, we present the first result that shows that natural body dynamics can be predicted at all. We built an LSTM network that is trained on violin and piano recital videos uploaded to the Internet. The predicted points are applied onto a rigged avatar to create the animation.
Link with videos https://arviolin.github.io/AudioBodyDynamics/
Cited by in corpus (11)
- Action2Motion: Conditioned Generation of 3D Human Motions
- Modeling Human Motion with Quaternion-based Neural Networks
- Rhythm is a Dancer: Music-Driven Motion Synthesis with Global Structure
- Temporally Guided Music-to-Body-Movement Generation
- DiffDance: Cascaded Human Motion Diffusion Model for Dance Generation
- DanceIt: Music-inspired Dancing Video Synthesis
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
- Synchronize Dual Hands for Physics-Based Dexterous Guitar Playing
- A Human-Computer Duet System for Music Performance
- ELGAR: Expressive Cello Performance Motion Generation for Audio Rendition