Survey on Visual Signal Coding and Processing with Generative Models: Technologies, Standards and Optimization
arXiv:2405.14221 · doi:10.1109/JETCAS.2024.3403524
Abstract
This paper provides a survey of the latest developments in visual signal coding and processing with generative models. Specifically, our focus is on presenting the advancement of generative models and their influence on research in the domain of visual signal coding and processing. This survey study begins with a brief introduction of well-established generative models, including the Variational Autoencoder (VAE) models, Generative Adversarial Network (GAN) models, Autoregressive (AR) models, Normalizing Flows and Diffusion models. The subsequent section of the paper explores the advancements in visual signal coding based on generative models, as well as the ongoing international standardization activities. In the realm of visual signal processing, our focus lies on the application and development of various generative models in the research of visual signal restoration. We also present the latest developments in generative visual signal synthesis and editing, along with visual signal quality assessment using generative models and quality assessment for generative models. The practical implementation of these studies is closely linked to the investigation of fast optimization. This paper additionally presents the latest advancements in fast optimization on visual signal coding and processing with generative models. We hope to advance this field by providing researchers and practitioners a comprehensive literature review on the topic of visual signal coding and processing with generative models.
References in corpus (47)
- Conditional Generative Adversarial Nets
- Diffusion Models in Vision: A Survey
- Variational image compression with a scale hyperprior
- End-to-end Optimized Image Compression
- Normalizing Flows: An Introduction and Review of Current Methods
- ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks
- The Perception-Distortion Tradeoff
- Artificial Intelligence in the Creative Industries: A Review
- CT Super-resolution GAN Constrained by the Identical, Residual, and Cycle Learning Ensemble(GAN-CIRCLE)
- Video Compression With Rate-Distortion Autoencoders
- M-LVC: Multiple Frames Prediction for Learned Video Compression
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression
- Causal Contextual Prediction for Learned Image Compression
- Generative Adversarial Networks and Perceptual Losses for Video Super-Resolution
- Learning for Video Compression
- Learning for Video Compression with Recurrent Auto-Encoder and Recurrent Probability Model
- Semantic Image Synthesis via Diffusion Models
- LDMVFI: Video Frame Interpolation with Latent Diffusion Models
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
- Advancing Learned Video Compression with In-loop Frame Prediction
- Towards Robust Neural Image Compression: Adversarial Attack and Model Finetuning
- Diffusion Models for Image Restoration and Enhancement: A Comprehensive Survey
- IDF++: Analyzing and Improving Integer Discrete Flows for Lossless Compression
- Frame Interpolation with Multi-Scale Deep Loss Functions and Generative Adversarial Networks
- Learning Cross-Scale Weighted Prediction for Efficient Neural Video Compression
- Overfitting for Fun and Profit: Instance-Adaptive Data Compression
- Quality Prediction on Deep Generative Images
- COIN++: Neural Compression Across Modalities
- EVC: Towards Real-Time Neural Image Compression with Mask Decay
- A Study on the Evaluation of Generative Models
- Lossy Compression with Gaussian Diffusion
- Iterative training of neural networks for intra prediction
- A Residual Diffusion Model for High Perceptual Quality Codec Augmentation
- High-Fidelity Image Compression with Score-based Generative Models
- Perceptually-inspired super-resolution of compressed videos
- Instance-Adaptive Video Compression: Improving Neural Codecs by Training on the Test Set
- Interactive Face Video Coding: A Generative Compression Framework
- Post-Training Quantization for Cross-Platform Learned Image Compression
- Compound Frechet Inception Distance for Quality Assessment of GAN Created Images
- Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration Information
- Accelerating Learnt Video Codecs with Gradient Decay and Layer-wise Distillation
- On the Choice of Perception Loss Function for Learned Video Compression
- Diffusion Models with Deterministic Normalizing Flow Priors
- Attack and Defense Analysis of Learned Image Compression
- Bandwidth-efficient Inference for Neural Image Compression
- A Training-Free Defense Framework for Robust Learned Image Compression