papers

Publications (70)

cs.CV2020

Learning to Blindly Assess Image Quality in the Laboratory and Wild

Weixia Zhang, Kede Ma, Guangtao Zhai +1

Computational models for blind image quality assessment (BIQA) are typically trained in well-controlled laboratory environments with limited generalizability to realistically disto…

cs.CV2024

Arbitrary-Scale Video Super-Resolution with Structural and Textural Priors

Wei Shang, Dongwei Ren, Wanying Zhang +3

Arbitrary-scale video super-resolution (AVSR) aims to enhance the resolution of video frames, potentially at various scaling factors, which presents several challenges regarding sp…

eess.IV2024

Learned Scanpaths Aid Blind Panoramic Video Quality Assessment

Kanglong Fan, Wen Wen, Mu Li +2

Panoramic videos have the advantage of providing an immersive and interactive viewing experience. Nevertheless, their spherical nature gives rise to various and uncertain user view…

cs.CV2021

Perceptually Optimized Deep High-Dynamic-Range Image Tone Mapping

Chenyang Le, Jiebin Yan, Yuming Fang +1

We describe a deep high-dynamic-range (HDR) image tone mapping operator that is computationally efficient and perceptually optimized. We first decompose an HDR image into a normali…

cs.CV2025

VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank

Tianhe Wu, Jian Zou, Jie Liang +2

DeepSeek-R1 has demonstrated remarkable effectiveness in incentivizing reasoning and generalization capabilities of large language models (LLMs) through reinforcement learning. Nev…

cs.CV2023

IconShop: Text-Guided Vector Icon Synthesis with Autoregressive Transformers

Ronghuan Wu, Wanchao Su, Kede Ma +1

Scalable Vector Graphics (SVG) is a popular vector image format that offers good support for interactivity and animation. Despite its appealing characteristics, creating custom SVG…

cs.CV2023

Blind Image Quality Assessment via Vision-Language Correspondence: A Multitask Learning Perspective

Weixia Zhang, Guangtao Zhai, Ying Wei +2

We aim at advancing blind image quality assessment (BIQA), which predicts the human perception of image quality without any reference information. We develop a general and automate…

cs.CV2024

Multiscale Sliced Wasserstein Distances as Perceptual Color Difference Measures

Jiaqi He, Zhihua Wang, Leon Wang +4

Contemporary color difference (CD) measures for photographic images typically operate by comparing co-located pixels, patches in a ``perceptually uniform'' color space, or features…

cs.LG2025

CLDyB: Towards Dynamic Benchmarking for Continual Learning with Pre-trained Models

Shengzhuang Chen, Yikai Liao, Xiaoxiao Sun +2

The advent of the foundation model era has sparked significant research interest in leveraging pre-trained representations for continual learning (CL), yielding a series of top-per…

eess.IV2024

Perceptual Assessment and Optimization of HDR Image Rendering

Peibei Cao, Rafal K. Mantiuk, Kede Ma

High dynamic range (HDR) rendering has the ability to faithfully reproduce the wide luminance ranges in natural scenes, but how to accurately assess the rendering quality is relati…

cs.CV2021

Troubleshooting Blind Image Quality Models in the Wild

Zhihua Wang, Haotao Wang, Tianlong Chen +2

Recently, the group maximum differentiation competition (gMAD) has been used to improve blind image quality assessment (BIQA) models, with the help of full-reference metrics. When…

cs.CV2019

dipIQ: Blind Image Quality Assessment by Learning-to-Rank Discriminable Image Pairs

Kede Ma, Wentao Liu, Tongliang Liu +2

Objective assessment of image quality is fundamentally important in many image processing tasks. In this work, we focus on learning blind image quality assessment (BIQA) models whi…

cs.LG2026

One-Token Verification for Reasoning Correctness Estimation

Zhan Zhuang, Xiequn Wang, Zebin Chen +4

Recent breakthroughs in large language models (LLMs) have led to notable successes in complex reasoning tasks, such as mathematical problem solving. A common strategy for improving…

cs.MM2019

Intrinsic Image Popularity Assessment

Keyan Ding, Kede Ma, Shiqi Wang

The goal of research in automatic image popularity assessment (IPA) is to develop computational models that can accurately predict the potential of a social image to go viral on th…

eess.IV2020

Efficient and Effective Context-Based Convolutional Entropy Modeling for Image Compression

Mu Li, Kede Ma, Jane You +2

Precise estimation of the probabilistic structure of natural images plays an essential role in image compression. Despite the recent remarkable success of end-to-end optimized imag…

cs.CV2026

RUTA: Principled Visual Token Allocation via Rate-Utility Optimization

Jian Zou, Xiaoyu Xu, Zhihua Wang +3

High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token se…

cs.LG2020

I Am Going MAD: Maximum Discrepancy Competition for Comparing Classifiers Adaptively

Haotao Wang, Tianlong Chen, Zhangyang Wang +1

The learning of hierarchical representations for image classification has experienced an impressive series of successes due in part to the availability of large-scale labeled data…

cs.CV2025

When No-Reference Image Quality Models Meet MAP Estimation in Diffusion Latents

Weixia Zhang, Dingquan Li, Guangtao Zhai +2

Contemporary no-reference image quality assessment (NR-IQA) models can effectively quantify perceived image quality, often achieving strong correlations with human perceptual score…

cs.CV2026

Learning Stable Canonical Worlds for Novel View Synthesis and Beyond

Xiaoyu Xu, Jian Zou, Sheyang Tang +3

Feed-forward Gaussian splatting (FFGS) facilitates real-time novel view synthesis, yet current methods often remain tied to view-dependent predictions. As more input views are adde…

cs.CV2024

Analysis of Video Quality Datasets via Design of Minimalistic Video Quality Models

Wei Sun, Wen Wen, Xiongkuo Min +3

Blind video quality assessment (BVQA) plays an indispensable role in monitoring and improving the end-users' viewing experience in various real-world video-enabled media applicatio…

cs.CV2025

Semantic Contextualization of Face Forgery: A New Definition, Dataset, and Detection Method

Mian Zou, Baosheng Yu, Yibing Zhan +2

In recent years, deep learning has greatly streamlined the process of manipulating photographic face images. Aware of the potential dangers, researchers have developed various tool…

cs.LG2025

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

Kehua Feng, Keyan Ding, Hongzhi Tan +8

Reliable evaluation of large language models (LLMs) is impeded by two key challenges: objective metrics often fail to reflect human perception of natural language, and exhaustive h…

cs.CV2026

RELO: Reinforcement Learning to Localize for Visual Object Tracking

Xin Chen, Chuanyu Sun, Jiao Xu +4

Conventional visual object trackers localize targets using handcrafted spatial priors, often in the form of heatmaps. Such priors provide only surrogate supervision and are poorly…

eess.IV2021

Pseudocylindrical Convolutions for Learned Omnidirectional Image Compression

Mu Li, Kede Ma, Jinxing Li +1

Although equirectangular projection (ERP) is a convenient form to store omnidirectional images (also known as 360-degree images), it is neither equal-area nor conformal, thus not f…

cs.CV2023

Image Quality Assessment: Integrating Model-Centric and Data-Centric Approaches

Peibei Cao, Dingquan Li, Kede Ma

Learning-based image quality assessment (IQA) has made remarkable progress in the past decade, but nearly all consider the two key components -- model and data -- in isolation. Spe…

cs.CV2018

Deep Blur Mapping: Exploiting High-Level Semantics by Deep Neural Networks

Kede Ma, Huan Fu, Tongliang Liu +2

The human visual system excels at detecting local blur of visual images, but the underlying mechanism is not well understood. Traditional views of blur such as reduction in energy…

cs.CV2021

Image Quality Assessment in the Modern Age

Kede Ma, Yuming Fang

This tutorial provides the audience with the basic theories, methodologies, and current progresses of image quality assessment (IQA). From an actionable perspective, we will first…

eess.IV2021

Locally Adaptive Structure and Texture Similarity for Image Quality Assessment

Keyan Ding, Yi Liu, Xueyi Zou +2

The latest advances in full-reference image quality assessment (IQA) involve unifying structure and texture similarity based on deep representations. The resulting Deep Image Struc…

eess.IV2021

Active Fine-Tuning from gMAD Examples Improves Blind Image Quality Assessment

Zhihua Wang, Kede Ma

The research in image quality assessment (IQA) has a long history, and significant progress has been made by leveraging recent advances in deep neural networks (DNNs). Despite high…

cs.CV2026

Diversity-Preserved Distribution Matching Distillation for Fast Visual Synthesis

Tianhe Wu, Ruibin Li, Lei Zhang +1

Distribution matching distillation (DMD) facilitates few-step image generation by aligning a distilled student with a reference multi-step teacher. In practice, however, optimizing…

cs.CV2025

Toward Generalized Image Quality Assessment: Relaxing the Perfect Reference Quality Assumption

Du Chen, Tianhe Wu, Kede Ma +1

Full-reference image quality assessment (FR-IQA) generally assumes that reference images are of perfect quality. However, this assumption is flawed due to the sensor and optical li…

cs.CV2021

Uncertainty-Aware Blind Image Quality Assessment in the Laboratory and Wild

Weixia Zhang, Kede Ma, Guangtao Zhai +1

Performance of blind image quality assessment (BIQA) models has been significantly boosted by end-to-end optimization of feature engineering and quality regression. Nevertheless, d…

cs.CV2026

Self-Supervised AI-Generated Image Detection: A Camera Metadata Perspective

Nan Zhong, Mian Zou, Yiran Xu +4

The proliferation of AI-generated imagery poses escalating challenges for multimedia forensics, yet many existing detectors depend on assumptions about the internals of specific ge…

cs.CV2025

Hiding Images in Diffusion Models by Editing Learned Score Functions

Haoyu Chen, Yunqiao Yang, Nan Zhong +1

Hiding data using neural networks (i.e., neural steganography) has achieved remarkable success across both discriminative classifiers and generative adversarial networks. However,…

cs.CV2023

Joint Video Multi-Frame Interpolation and Deblurring under Unknown Exposure Time

Wei Shang, Dongwei Ren, Yi Yang +3

Natural videos captured by consumer cameras often suffer from low framerate and motion blur due to the combination of dynamic scene complexity, lens and sensor imperfection, and le…

cs.LG2025

Dataset Distillation as Data Compression: A Rate-Utility Perspective

Youneng Bao, Yiping Liu, Zhuo Chen +3

Driven by the ``scale-is-everything'' paradigm, modern machine learning increasingly demands ever-larger datasets and models, yielding prohibitive computational and storage require…

eess.IV2019

Blind Image Quality Assessment Using A Deep Bilinear Convolutional Neural Network

Weixia Zhang, Kede Ma, Jia Yan +2

We propose a deep bilinear model for blind image quality assessment (BIQA) that handles both synthetic and authentic distortions. Our model consists of two convolutional neural net…

cs.LG2025

SD-LoRA: Scalable Decoupled Low-Rank Adaptation for Class Incremental Learning

Yichen Wu, Hongming Piao, Long-Kai Huang +6

Continual Learning (CL) with foundation models has recently emerged as a promising paradigm to exploit abundant knowledge acquired during pre-training for tackling sequential tasks…

eess.IV2024

Modular Blind Video Quality Assessment

Wen Wen, Mu Li, Yabin Zhang +4

Blind video quality assessment (BVQA) plays a pivotal role in evaluating and improving the viewing experience of end-users across a wide range of video-based platforms and services…

cs.CV2026

MDS-VQA: Model-Informed Data Selection for Video Quality Assessment

Jian Zou, Xiaoyu Xu, Zhihua Wang +3

Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centri…

eess.IV2021

Perceptual Quality Assessment of Omnidirectional Images as Moving Camera Videos

Xiangjie Sui, Kede Ma, Yiru Yao +1

Omnidirectional images (also referred to as static 360° panoramas) impose viewing conditions much different from those of regular 2D images. How do humans perceive image distortio…

cs.CV2023

Learning a Deep Color Difference Metric for Photographic Images

Haoyu Chen, Zhihua Wang, Yang Yang +2

Most well-established and widely used color difference (CD) metrics are handcrafted and subject-calibrated against uniformly colored patches, which do not generalize well to photog…

cs.CV2024

AniClipart: Clipart Animation with Text-to-Video Priors

Ronghuan Wu, Wanchao Su, Kede Ma +1

Clipart, a pre-made art form, offers a convenient and efficient way of creating visual content. However, traditional workflows for animating static clipart are laborious and time-c…

cs.CV2026

PrISM-IQA: Image Quality Assessment Made Practical for Smartphone Photography

Shuyan Zhai, Jiaqi He, Weixia Zhang +4

Existing smartphone image quality assessment (IQA) methods commonly reduce perceptual quality to a single score. However, this scalar formulation is poorly aligned with practical i…

cs.CV2026

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

Ce Chen, Congrui Wang, Yonglin Li +23

Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant challenges for edge deployment in terms of l…

cs.CV2026

X2HDR: HDR Image Generation in a Perceptually Uniform Space

Ronghuan Wu, Wanchao Su, Kede Ma +2

High-dynamic-range (HDR) formats and displays are becoming increasingly prevalent, yet state-of-the-art image generators (e.g., Stable Diffusion and FLUX) typically remain limited…

cs.CV2025

Scanpath Prediction in Panoramic Videos via Expected Code Length Minimization

Mu Li, Kanglong Fan, Kede Ma

Predicting human scanpaths when exploring panoramic videos is a challenging task due to the spherical geometry and the multimodality of the input, and the inherent uncertainty and…

eess.IV2024

Steerable Pyramid Transform Enables Robust Left Ventricle Quantification

Xiangyang Zhu, Kede Ma, Wufeng Xue

Predicting cardiac indices has long been a focal point in the medical imaging community. While various deep learning models have demonstrated success in quantifying cardiac indices…

cs.CV2021

Semi-Supervised Deep Ensembles for Blind Image Quality Assessment

Zhihua Wang, Dingquan Li, Kede Ma

Ensemble methods are generally regarded to be better than a single model if the base learners are deemed to be "accurate" and "diverse." Here we investigate a semi-supervised ensem…

cs.CV2024

Learning Where to Edit Vision Transformers

Yunqiao Yang, Long-Kai Huang, Shengzhuang Chen +2

Model editing aims to data-efficiently correct predictive errors of large pre-trained models while ensuring generalization to neighboring failures and locality to minimize unintend…

cs.CV2024

Perceptual Quality Assessment of Virtual Reality Videos in the Wild

Wen Wen, Mu Li, Yiru Yao +5

Investigating how people perceive virtual reality (VR) videos in the wild (i.e., those captured by everyday users) is a crucial and challenging task in VR-related applications due…

eess.IV2020

Characterizing Generalized Rate-Distortion Performance of Video Coding: An Eigen Analysis Approach

Zhengfang Duanmu, Wentao Liu, Zhuoran Li +2

Rate-distortion (RD) theory is at the heart of lossy data compression. Here we aim to model the generalized RD (GRD) trade-off between the visual quality of a compressed video and…

cs.CV2025

Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding Approach

Mian Zou, Baosheng Yu, Yibing Zhan +2

In recent years, the multimedia forensics and security community has seen remarkable progress in multitask learning for DeepFake (i.e., face forgery) detection. The prevailing appr…

cs.CV2025

Bi-Level Optimization for Self-Supervised AI-Generated Face Detection

Mian Zou, Nan Zhong, Baosheng Yu +2

AI-generated face detectors trained via supervised learning typically rely on synthesized images from specific generators, limiting their generalization to emerging generative tech…

eess.IV2023

A Perceptually Optimized and Self-Calibrated Tone Mapping Operator

Peibei Cao, Chenyang Le, Yuming Fang +1

With the increasing popularity and accessibility of high dynamic range (HDR) photography, tone mapping operators (TMOs) for dynamic range compression are practically demanding. In…

cs.MM2026

Subjective Evaluation of Frame Rate in Bitrate-Constrained Live Streaming

Jiaqi He, Zhengfang Duanmu, Kede Ma

Bandwidth constraints in live streaming require video codecs to balance compression strength and frame rate, yet the perceptual consequences of this trade-off remain underexplored.…

cs.CV2024

A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment

Tianhe Wu, Kede Ma, Jie Liang +2

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, inter…

cs.CV2025

Self-Supervised Learning for Detecting AI-Generated Faces as Anomalies

Mian Zou, Baosheng Yu, Yibing Zhan +1

The detection of AI-generated faces is commonly approached as a binary classification task. Nevertheless, the resulting detectors frequently struggle to adapt to novel AI face gene…

eess.IV2024

Learned HDR Image Compression for Perceptually Optimal Storage and Display

Peibei Cao, Haoyu Chen, Jingzhe Ma +5

High dynamic range (HDR) capture and display have seen significant growth in popularity driven by the advancements in technology and increasing consumer demand for superior image q…

cs.CV2022

Perceptual Attacks of No-Reference Image Quality Models with Human-in-the-Loop

Weixia Zhang, Dingquan Li, Xiongkuo Min +4

No-reference image quality assessment (NR-IQA) aims to quantify how humans perceive visual distortions of digital images without access to their undistorted references. NR-IQA mode…

cs.CV2022

Continual Learning for Blind Image Quality Assessment

Weixia Zhang, Dingquan Li, Chao Ma +3

The explosive growth of image data facilitates the fast development of image processing and computer vision methods for emerging visual applications, meanwhile introducing novel di…

eess.IV2024

Hierarchical Prior-based Super Resolution for Point Cloud Geometry Compression

Dingquan Li, Kede Ma, Jing Wang +1

The Geometry-based Point Cloud Compression (G-PCC) has been developed by the Moving Picture Experts Group to compress point clouds. In its lossy mode, the reconstructed point cloud…

eess.IV2020

Comparison of Image Quality Models for Optimization of Image Processing Systems

Keyan Ding, Kede Ma, Shiqi Wang +1

The performance of objective image quality assessment (IQA) models has been evaluated primarily by comparing model predictions to human quality judgments. Perceptual datasets gathe…

cs.CV2020

Image Quality Assessment: Unifying Structure and Texture Similarity

Keyan Ding, Kede Ma, Shiqi Wang +1

Objective measures of image quality generally operate by comparing pixels of a "degraded" image to those of the original. Relative to human observers, these measures are overly sen…

cs.CR2022

Hiding Images in Deep Probabilistic Models

Haoyu Chen, Linqi Song, Zhenxing Qian +2

Data hiding with deep neural networks (DNNs) has experienced impressive successes in recent years. A prevailing scheme is to train an autoencoder, consisting of an encoding network…

cs.LG2025

Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation

Zhan Zhuang, Xiequn Wang, Wei Li +9

Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal mini…

cs.CV2023

Measuring Perceptual Color Differences of Smartphone Photographs

Zhihua Wang, Keshuo Xu, Yang Yang +5

Measuring perceptual color differences (CDs) is of great importance in modern smartphone photography. Despite the long history, most CD measures have been constrained by psychophys…

cs.CV2025

NTIRE 2025 Challenge on Real-World Face Restoration: Methods and Results

Zheng Chen, Jingkai Wang, Kai Liu +51

This paper provides a review of the NTIRE 2025 challenge on real-world face restoration, highlighting the proposed solutions and the resulting outcomes. The challenge focuses on ge…

cs.CV2024

Task-Specific Normalization for Continual Learning of Blind Image Quality Models

Weixia Zhang, Kede Ma, Guangtao Zhai +1

In this paper, we present a simple yet effective continual learning method for blind image quality assessment (BIQA) with improved quality prediction accuracy, plasticity-stability…

cs.CV2021

Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition

Jiebin Yan, Yu Zhong, Yuming Fang +2

Semantic segmentation is an extensively studied task in computer vision, with numerous methods proposed every year. Thanks to the advent of deep learning in semantic segmentation,…