Deep Generative Video Compression
arXiv:1810.02845
Abstract
The usage of deep generative models for image compression has led to impressive performance gains over classical codecs while neural video compression is still in its infancy. Here, we propose an end-to-end, deep generative modeling approach to compress temporal sequences with a focus on video. Our approach builds upon variational autoencoder (VAE) models for sequential data and combines them with recent work on neural image compression. The approach jointly learns to transform the original sequence into a lower-dimensional representation as well as to discretize and entropy code this representation according to predictions of the sequential VAE. Rate-distortion evaluations on small videos from public data sets with varying complexity and diversity show that our model yields competitive results when trained on generic video content. Extreme compression performance is achieved when training the model on specialized content.
Accepted at NeurIPS 2019, 15 pages, 8 figures
References in corpus (15)
- Auto-Encoding Variational Bayes
- The Kinetics Human Action Video Dataset
- Variational image compression with a scale hyperprior
- End-to-end Optimized Image Compression
- Variational Inference with Normalizing Flows
- A Recurrent Latent Variable Model for Sequential Data
- Soft-to-Hard Vector Quantization for End-to-End Learning Compressible Representations
- Joint Autoregressive and Hierarchical Priors for Learned Image Compression
- Stochastic Video Generation with a Learned Prior
- Disentangling factors of variation in deep representations using adversarial training
- Video Compression With Rate-Distortion Autoencoders
- Stochastic Adversarial Video Prediction
- Learning for Video Compression
- Deep Kalman Filters
- Variational Tempering