The Perception-Distortion Tradeoff
arXiv:1711.06077 · doi:10.1109/CVPR.2018.00652
Abstract
Image restoration algorithms are typically evaluated by some distortion measure (e.g. PSNR, SSIM, IFC, VIF) or by human opinion scores that quantify perceived perceptual quality. In this paper, we prove mathematically that distortion and perceptual quality are at odds with each other. Specifically, we study the optimal probability for correctly discriminating the outputs of an image restoration algorithm from real images. We show that as the mean distortion decreases, this probability must increase (indicating worse perceptual quality). As opposed to the common belief, this result holds true for any distortion measure, and is not only a problem of the PSNR or SSIM criteria. We also show that generative-adversarial-nets (GANs) provide a principled way to approach the perception-distortion bound. This constitutes theoretical support to their observed success in low-level vision tasks. Based on our analysis, we propose a new methodology for evaluating image restoration methods, and use it to perform an extensive comparison between recent super-resolution algorithms.
CVPR 2018 (long oral presentation), see talk at: https://youtu.be/_aXbGqdEkjk?t=39m43s
References in corpus (4)
Cited by in corpus (152)
- ESRGAN: Enhanced Super-Resolution Generative Adversarial Networks
- Deep Learning for Single Image Super-Resolution: A Brief Review
- The Perception-Distortion Tradeoff
- Unsupervised Single Image Dehazing Using Dark Channel Prior Loss
- Point2Mesh: A Self-Prior for Deformable Meshes
- Comparison of Image Quality Models for Optimization of Image Processing Systems
- Learning Spatial Attention for Face Super-Resolution
- Learning Temporal Coherence via Self-Supervision for GAN-based Video Generation
- Deep Learning-Based Video Coding: A Review and A Case Study
- Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement
- Large Scale Image Completion via Co-Modulated Generative Adversarial Networks
- Face Generation and Editing with StyleGAN: A Survey
- Real-ESRGAN: Training Real-World Blind Super-Resolution with Pure Synthetic Data
- Deep Learning for Image Super-resolution: A Survey
- ProxIQA: A Proxy Approach to Perceptual Optimization of Learned Image Compression
- Bridging the Gap Between Computational Photography and Visual Recognition
- Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff
- A Latent Encoder Coupled Generative Adversarial Network (LE-GAN) for Efficient Hyperspectral Image Super-resolution
- Edge and Identity Preserving Network for Face Super-Resolution
- ST-MFNet: A Spatio-Temporal Multi-Flow Network for Frame Interpolation
- Two-Stream Action Recognition-Oriented Video Super-Resolution
- Optimal Transport for Unsupervised Denoising Learning
- Generative Adversarial Networks for Image and Video Synthesis: Algorithms and Applications
- Benchmarking Deep Learning-Based Low-Dose CT Image Denoising Algorithms
- Stimulating Diffusion Model for Image Denoising via Adaptive Embedding and Ensembling
- High Perceptual Quality Image Denoising with a Posterior Sampling CGAN
- Image Quality Assessment for Perceptual Image Restoration: A New Dataset, Benchmark and Metric
- A Systematic Survey of Deep Learning-based Single-Image Super-Resolution
- Rate-Distortion-Perception Tradeoff of Variable-Length Source Coding for General Information Sources
- Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
- PyNET-CA: Enhanced PyNET with Channel Attention for End-to-End Mobile Image Signal Processing
- Hierarchical Quantized Autoencoders
- On The Classification-Distortion-Perception Tradeoff
- Universal Rate-Distortion-Perception Representations for Lossy Compression
- Enhanced Standard Compatible Image Compression Framework based on Auxiliary Codec Networks
- Uncovering the Over-smoothing Challenge in Image Super-Resolution: Entropy-based Quantification and Contrastive Optimization
- Towards Realistic Face Photo-Sketch Synthesis via Composition-Aided GANs
- FreeStyleGAN: Free-view Editable Portrait Rendering with the Camera Manifold
- Controlling Rate, Distortion, and Realism: Towards a Single Comprehensive Neural Image Compression Model
- (ASNA) An Attention-based Siamese-Difference Neural Network with Surrogate Ranking Loss function for Perceptual Image Quality Assessment
- DriftRec: Adapting diffusion models to blind JPEG restoration
- TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution
- IFQA: Interpretable Face Quality Assessment
- Waveforms for Computing Over the Air
- Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures
- Distributed Learning and Inference with Compressed Images
- A deep cascade of ensemble of dual domain networks with gradient-based T1 assistance and perceptual refinement for fast MRI reconstruction
- Towards Effective and Interpretable Semantic Communications
- On the advantages of stochastic encoders
- Survey on Visual Signal Coding and Processing with Generative Models: Technologies, Standards and Optimization
- Context-Aware Image Matting for Simultaneous Foreground and Alpha Estimation
- PIPAL: a Large-Scale Image Quality Assessment Dataset for Perceptual Image Restoration
- Instrument-To-Instrument translation: Instrumental advances drive restoration of solar observation series via deep learning
- Attention-Based Generative Neural Image Compression on Solar Dynamics Observatory
- Wavelet Domain Style Transfer for an Effective Perception-distortion Tradeoff in Single Image Super-Resolution
- Creating High Resolution Images with a Latent Adversarial Generator
- Perception Consistency Ultrasound Image Super-resolution via Self-supervised CycleGAN
- Generating Diverse Structure for Image Inpainting With Hierarchical VQ-VAE
- NAWQ-SR: A Hybrid-Precision NPU Engine for Efficient On-Device Super-Resolution
- Diffusion Posterior Proximal Sampling for Image Restoration
- On the Importance of Denoising when Learning to Compress Images
- GAN- vs. JPEG2000 Image Compression for Distributed Automotive Perception: Higher Peak SNR Does Not Mean Better Semantic Segmentation
- Multi-scale Processing of Noisy Images using Edge Preservation Losses
- Camera Lens Super-Resolution
- Variational Bayes image restoration with compressive autoencoders
- WaveFill: A Wavelet-based Generation Network for Image Inpainting
- MDCN: Multi-scale Dense Cross Network for Image Super-Resolution
- Neural-based Compression Scheme for Solar Image Data
- Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based Approach
- A Compression Objective and a Cycle Loss for Neural Image Compression
- Validation and Generalizability of Self-Supervised Image Reconstruction Methods for Undersampled MRI
- A Theory of the Distortion-Perception Tradeoff in Wasserstein Space
- SNIPS: Solving Noisy Inverse Problems Stochastically
- On Perceptual Lossy Compression: The Cost of Perceptual Reconstruction and An Optimal Training Framework
- DiVa-360: The Dynamic Visual Dataset for Immersive Neural Fields
- Natural and Realistic Single Image Super-Resolution with Explicit Natural Manifold Discrimination
- AIM 2020 Challenge on Video Extreme Super-Resolution: Methods and Results
- Using the Semantic Information G Measure to Explain and Extend Rate-Distortion Functions and Maximum Entropy Distributions
- Perceptual-Distortion Balanced Image Super-Resolution is a Multi-Objective Optimization Problem
- Image Resizing by Reconstruction from Deep Features
- CFSNet: Toward a Controllable Feature Space for Image Restoration
- Learning to Zoom-in via Learning to Zoom-out: Real-world Super-resolution by Generating and Adapting Degradation
- See360: Novel Panoramic View Interpolation
- EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution
- Elevating 3D Models: High-Quality Texture and Geometry Refinement from a Low-Quality Model
- Toward Bridging the Simulated-to-Real Gap: Benchmarking Super-Resolution on Real Data
- Real-Time Super-Resolution System of 4K-Video Based on Deep Learning
- Fine-grained Attention and Feature-sharing Generative Adversarial Networks for Single Image Super-Resolution
- Learning End-to-End Lossy Image Compression: A Benchmark
- Multi-level Wavelet-based Generative Adversarial Network for Perceptual Quality Enhancement of Compressed Video
- WiSoSuper: Benchmarking Super-Resolution Methods on Wind and Solar Data
- Wavelet-Based Dual-Branch Network for Image Demoireing
- Reference-Based Face Super-Resolution Using the Spatial Transformer
- A coding theorem for the rate-distortion-perception function
- Scale-Equivariant Imaging: Self-Supervised Learning for Image Super-Resolution and Deblurring
- Analyzing α-divergence in Gaussian Rate-Distortion-Perception Theory
- From Rank Estimation to Rank Approximation: Rank Residual Constraint for Image Restoration
- Variational AutoEncoder for Reference based Image Super-Resolution
- A statistically constrained internal method for single image super-resolution
- Rethinking Deep Image Prior for Denoising
- Posterior Sampling for Image Restoration using Explicit Patch Priors
- Learning to Enhance Low-Light Image via Zero-Reference Deep Curve Estimation
- DiffFuSR: Super-Resolution of all Sentinel-2 Multispectral Bands using Diffusion Models
- Progressively Unfreezing Perceptual GAN
- Smoother Network Tuning and Interpolation for Continuous-level Image Processing
- SRZoo: An integrated repository for super-resolution using deep learning
- A Regularized Conditional GAN for Posterior Sampling in Image Recovery Problems
- Perceptual Image Super-Resolution with Progressive Adversarial Network
- RankSRGAN: Super Resolution Generative Adversarial Networks with Learning to Rank
- Multi-Scale Texture Loss for CT denoising with GANs
- Advancing Limited-Angle CT Reconstruction Through Diffusion-Based Sinogram Completion
- Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
- Implicit Subspace Prior Learning for Dual-Blind Face Restoration
- Projected Latent Markov Chain Monte Carlo: Conditional Sampling of Normalizing Flows
- Is There Tradeoff between Spatial and Temporal in Video Super-Resolution?
- Softmax Splatting for Video Frame Interpolation
- Conditional Adversarial Camera Model Anonymization
- Saliency Driven Perceptual Image Compression
- HRINet: Alternative Supervision Network for High-resolution CT image Interpolation
- DR-KFS: A Differentiable Visual Similarity Metric for 3D Shape Reconstruction
- End-to-End Image Compression with Probabilistic Decoding
- Long-Term Human Video Generation of Multiple Futures Using Poses
- Deep Likelihood Network for Image Restoration with Multiple Degradation Levels
- A HVS-inspired Attention to Improve Loss Metrics for CNN-based Perception-Oriented Super-Resolution
- Region-Adaptive Deformable Network for Image Quality Assessment
- Deep Multiple Description Coding by Learning Scalar Quantization
- Multi-Scale Recursive and Perception-Distortion Controllable Image Super-Resolution
- The Unreasonable Effectiveness of Texture Transfer for Single Image Super-resolution
- Distribution Preserving Source Separation With Time Frequency Predictive Models
- Deep 3D Pan via adaptive "t-shaped" convolutions with global and local adaptive dilations
- Deep Optimized Multiple Description Image Coding via Scalar Quantization Learning
- Contrastive Feature Loss for Image Prediction
- Editorial: Introduction to the Issue on Deep Learning for Image/Video Restoration and Compression
- Learning Deep Image Priors for Blind Image Denoising
- GRay: Ray Tracing 3D Gaussians Near the Speed of Splats
- Image restoration quality assessment based on regional differential information entropy
- GIFnets: Differentiable GIF Encoding Framework
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion
- Analyzing Perception-Distortion Tradeoff using Enhanced Perceptual Super-resolution Network
- A Novel Image Similarity Metric for Scene Composition Structure
- Bi-GANs-ST for Perceptual Image Super-resolution
- Infusion: internal diffusion for inpainting of dynamic textures and complex motion
- Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution
- Regularized Adaptation for Stable and Efficient Continuous-Level Learning on Image Processing Networks
- Joint Demosaicing and Super-Resolution (JDSR): Network Design and Perceptual Optimization
- Generative adversarial network-based image super-resolution using perceptual content losses
- Image Super-Resolution using Explicit Perceptual Loss
- Customized OCT images compression scheme with deep neural network
- Multi-Grid Back-Projection Networks
- Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
- Neural Enhancement in Content Delivery Systems: The State-of-the-Art and Future Directions
- Improving the Perceptual Quality of 2D Animation Interpolation