Disentangled Sequential Autoencoder
arXiv:1803.02991
Abstract
We present a VAE architecture for encoding and generating high dimensional sequential data, such as video or audio. Our deep generative model learns a latent representation of the data which is split into a static and dynamic part, allowing us to approximately disentangle latent time-dependent features (dynamics) from features which are preserved over time (content). This architecture gives us partial control over generating content and dynamics by conditioning on either one of these sets of features. In our experiments on artificially generated cartoon video clips and voice recordings, we show that we can convert the content of a given sequence into another one by such content swapping. For audio, this allows us to convert a male speaker into a female speaker and vice versa, while for video we can separately manipulate shapes and dynamics. Furthermore, we give empirical evidence for the hypothesis that stochastic RNNs as latent state models are more efficient at compressing and generating long sequences than deterministic ones, which may be relevant for applications in video compression.
Cited by in corpus (31)
- Recent Advances in Autoencoder-Based Representation Learning
- Dynamical Variational Autoencoders: A Comprehensive Review
- Are Disentangled Representations Helpful for Abstract Visual Reasoning?
- Weakly-Supervised Disentanglement Without Compromises
- Disentangling Factors of Variation Using Few Labels
- A Causal View on Robustness of Neural Networks
- A multimodal dynamical variational autoencoder for audiovisual speech representation learning
- Disentangled State Space Representations
- Contrastively Disentangled Sequential Variational Autoencoder
- Music FaderNets: Controllable Music Generation Based On High-Level Features via Low-Level Feature Modelling
- PDE-Driven Spatiotemporal Disentanglement
- Bayesian Optimization in Variational Latent Spaces with Dynamic Compression
- SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition
- Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
- Product of Orthogonal Spheres Parameterization for Disentangled Representation Learning
- On the Fairness of Disentangled Representations
- Hierarchical Autoregressive Modeling for Neural Video Compression
- SurpriseNet: Melody Harmonization Conditioning on User-controlled Surprise Contours
- On Disentanglement in Gaussian Process Variational Autoencoders
- Improving Disentangled Text Representation Learning with Information-Theoretic Guidance
- Customizing Sequence Generation with Multi-Task Dynamical Systems
- Learning Independently-Obtainable Reward Functions
- Analytic Manifold Learning: Unifying and Evaluating Representations for Continuous Control
- Capturing Actionable Dynamics with Structured Latent Ordinary Differential Equations
- Mixture factorized auto-encoder for unsupervised hierarchical deep factorization of speech signal
- Unsupervised Learning of Neurosymbolic Encoders
- Variational inference formulation for a model-free simulation of a dynamical system with unknown parameters by a recurrent neural network
- Towards Better Understanding of Disentangled Representations via Mutual Information
- A Benchmark of Dynamical Variational Autoencoders applied to Speech Spectrogram Modeling
- A Variational Auto-Encoder Model for Stochastic Point Processes
- Disentangled Dynamic Representations from Unordered Data