BVI-DVC: A Training Database for Deep Video Compression
arXiv:2003.13552 · doi:10.1109/TMM.2021.3108943
Abstract
Deep learning methods are increasingly being applied in the optimisation of video compression algorithms and can achieve significantly enhanced coding gains, compared to conventional approaches. Such approaches often employ Convolutional Neural Networks (CNNs) which are trained on databases with relatively limited content coverage. In this paper, a new extensive and representative video database, BVI-DVC, is presented for training CNN-based video compression systems, with specific emphasis on machine learning tools that enhance conventional coding architectures, including spatial resolution and bit depth up-sampling, post-processing and in-loop filtering. BVI-DVC contains 800 sequences at various spatial resolutions from 270p to 2160p and has been evaluated on ten existing network architectures for four different coding tools. Experimental results show that this database produces significant improvements in terms of coding gains over three existing (commonly used) image/video training databases under the same training and evaluation configurations. The overall additional coding improvements by using the proposed database for all tested coding modules and CNN architectures are up to 10.3% based on the assessment of PSNR and 8.1% based on VMAF.
References in corpus (8)
- Image and Video Compression with Neural Networks: A Review
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- M-LVC: Multiple Frames Prediction for Learned Video Compression
- Do We Need More Training Data?
- MFRNet: A New CNN Architecture for Post-Processing and In-loop Filtering
- Video Compression with CNN-based Post Processing
- Perceptually-inspired super-resolution of compressed videos
- Video compression with low complexity CNN-based spatial resolution adaptation
Cited by in corpus (14)
- Temporal Context Mining for Learned Video Compression
- CVEGAN: A Perceptually-inspired GAN for Compressed Video Enhancement
- RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training
- BVI-CC: A Dataset for Research on Video Compression and Quality Assessment
- ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video Compression
- BVI-AOM: A New Training Dataset for Deep Video Compression Optimization
- Enhancing Deformable Convolution based Video Frame Interpolation with Coarse-to-fine 3D CNN
- Full-reference Video Quality Assessment for User Generated Content Transcoding
- RankDVQA-mini: Knowledge Distillation-Driven Deep Video Quality Assessment
- MVAD: A Multiple Visual Artifact Detector for Video Streaming
- RMT-BVQA: Recurrent Memory Transformer-based Blind Video Quality Assessment for Enhanced Video Content
- RTSR: A Real-Time Super-Resolution Model for AV1 Compressed Content
- DaBiT: Depth and Blur informed Transformer for Video Focal Deblurring
- Enhancing HDR Video Compression through CNN-based Effective Bit Depth Adaptation