Dynamical Variational Autoencoders: A Comprehensive Review
arXiv:2008.12595 · doi:10.1561/2200000089
Abstract
Variational autoencoders (VAEs) are powerful deep generative models widely used to represent high-dimensional complex data through a low-dimensional latent space learned in an unsupervised manner. In the original VAE model, the input data vectors are processed independently. Recently, a series of papers have presented different extensions of the VAE to process sequential data, which model not only the latent space but also the temporal dependencies within a sequence of data vectors and corresponding latent vectors, relying on recurrent neural networks or state-space models. In this paper, we perform a literature review of these models. We introduce and discuss a general class of models, called dynamical variational autoencoders (DVAEs), which encompasses a large subset of these temporal VAE extensions. Then, we present in detail seven recently proposed DVAE models, with an aim to homogenize the notations and presentation lines, as well as to relate these models with existing classical temporal models. We have reimplemented those seven DVAE models and present the results of an experimental benchmark conducted on the speech analysis-resynthesis task (the PyTorch code is made publicly available). The paper concludes with a discussion on important issues concerning the DVAE class of models and future research guidelines.
References in corpus (18)
- WaveNet: A Generative Model for Raw Audio
- Modeling Temporal Dependencies in High-Dimensional Sequences: Application to Polyphonic Music Generation and Transcription
- Decomposing Motion and Content for Natural Video Sequence Prediction
- NVAE: A Deep Hierarchical Variational Autoencoder
- Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
- Learning Stochastic Recurrent Networks
- Variational Recurrent Auto-Encoders
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
- Learning Disentangled Representations with Semi-Supervised Deep Generative Models
- Lagging Inference Networks and Posterior Collapse in Variational Autoencoders
- Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder
- A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
- Tensor Factorization via Matrix Factorization
- Recurrent switching linear dynamical systems
- Tackling Over-pruning in Variational Autoencoders
- The Usual Suspects? Reassessing Blame for VAE Posterior Collapse
- A Sober Look at the Unsupervised Learning of Disentangled Representations and their Evaluation
- Neural Granular Sound Synthesis
Cited by in corpus (11)
- Self-Supervised Speech Representation Learning: A Review
- A Survey of Sound Source Localization with Deep Learning Methods
- Speech Enhancement and Dereverberation with Diffusion-based Generative Models
- DANSE: Data-driven Non-linear State Estimation of Model-free Process in Unsupervised Learning Setup
- A multimodal dynamical variational autoencoder for audiovisual speech representation learning
- Searching for Anomalies in the ZTF Catalog of Periodic Variable Stars
- Learning and controlling the source-filter representation of speech with a variational autoencoder
- Deep learning and differential equations for modeling changes in individual-level latent dynamics between observation periods
- From reductionism to realism: Holistic mathematical modelling for complex biological systems
- Variational Autoencoder for Calibration: A New Approach
- Using matrix-product states for time-series machine learning