RankDVQA: Deep VQA based on Ranking-inspired Hybrid Training
arXiv:2202.08595 · doi:10.1109/WACV57701.2024.00167
Abstract
In recent years, deep learning techniques have shown significant potential for improving video quality assessment (VQA), achieving higher correlation with subjective opinions compared to conventional approaches. However, the development of deep VQA methods has been constrained by the limited availability of large-scale training databases and ineffective training methodologies. As a result, it is difficult for deep VQA approaches to achieve consistently superior performance and model generalization. In this context, this paper proposes new VQA methods based on a two-stage training methodology which motivates us to develop a large-scale VQA training database without employing human subjects to provide ground truth labels. This method was used to train a new transformer-based network architecture, exploiting quality ranking of different distorted sequences rather than minimizing the difference from the ground-truth quality labels. The resulting deep VQA methods (for both full reference and no reference scenarios), FR- and NR-RankDVQA, exhibit consistently higher correlation with perceptual quality compared to the state-of-the-art conventional and deep VQA methods, with average SROCC values of 0.8972 (FR) and 0.7791 (NR) over eight test sets without performing cross-validation. The source code of the proposed quality metrics and the large training database are available at https://chenfeng-bristol.github.io/RankDVQA.
8 pages, 5 figures accepted by WACV 2024
References in corpus (5)
- Blind Image Quality Assessment Using A Deep Bilinear Convolutional Neural Network
- dipIQ: Blind Image Quality Assessment by Learning-to-Rank Discriminable Image Pairs
- Image Quality Assessment using Contrastive Learning
- Unified Quality Assessment of In-the-Wild Videos with Mixed Datasets Training
- Perceptually-inspired super-resolution of compressed videos
Cited by in corpus (7)
- Advances in Artificial Intelligence: A Review for the Creative Industries
- BVI-Artefact: An Artefact Detection Benchmark Dataset for Streamed Videos
- Full-reference Video Quality Assessment for User Generated Content Transcoding
- RMT-BVQA: Recurrent Memory Transformer-based Blind Video Quality Assessment for Enhanced Video Content
- MVAD: A Multiple Visual Artifact Detector for Video Streaming
- RankDVQA-mini: Knowledge Distillation-Driven Deep Video Quality Assessment
- Causal Debiasing for Visual Commonsense Reasoning