M-LVC: Multiple Frames Prediction for Learned Video Compression
arXiv:2004.10290 · doi:10.1109/CVPR42600.2020.00360
Abstract
We propose an end-to-end learned video compression scheme for low-latency scenarios. Previous methods are limited in using the previous one frame as reference. Our method introduces the usage of the previous multiple frames as references. In our scheme, the motion vector (MV) field is calculated between the current frame and the previous one. With multiple reference frames and associated multiple MV fields, our designed network can generate more accurate prediction of the current frame, yielding less residual. Multiple reference frames also help generate MV prediction, which reduces the coding cost of MV field. We use two deep auto-encoders to compress the residual and the MV, respectively. To compensate for the compression error of the auto-encoders, we further design a MV refinement network and a residual refinement network, taking use of the multiple reference frames as well. All the modules in our scheme are jointly optimized through a single rate-distortion loss function. We use a step-by-step training strategy to optimize the entire scheme. Experimental results show that the proposed method outperforms the existing learned video compression methods for low-latency mode. Our method also performs better than H.265 in both PSNR and MS-SSIM. Our code and models are publicly available.
Accepted to appear in CVPR2020; camera-ready
Cited by in corpus (16)
- Temporal Context Mining for Learned Video Compression
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression
- BVI-DVC: A Training Database for Deep Video Compression
- Advancing Learned Video Compression with In-loop Frame Prediction
- Learning Cross-Scale Weighted Prediction for Efficient Neural Video Compression
- CVEGAN: A Perceptually-inspired GAN for Compressed Video Enhancement
- Exploring Long- and Short-Range Temporal Information for Learned Video Compression
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video Compression
- Survey on Visual Signal Coding and Processing with Generative Models: Technologies, Standards and Optimization
- Learned Scalable Video Coding For Humans and Machines
- Towards Real-Time Neural Video Codec for Cross-Platform Application Using Calibration Information
- Learning-Based Conditional Image Coder Using Color Separation
- LVC-LGMC: Joint Local and Global Motion Compensation for Learned Video Compression
- CodingHomo: Bootstrapping Deep Homography With Video Coding
- Content-Adaptive Inference for State-of-the-art Learned Video Compression
- Neural Video Compression with In-Loop Contextual Filtering and Out-of-Loop Reconstruction Enhancement