papers

Publications (57)

cs.CV2024

Latent Spatiotemporal Adaptation for Generalized Face Forgery Video Detection

Daichi Zhang, Zihao Xiao, Jianmin Li +1

Face forgery videos have caused severe public concerns, and many detectors have been proposed. However, most of these detectors suffer from limited generalization when detecting vi…

cs.CV2024

Distilling Channels for Efficient Deep Tracking

Shiming Ge, Zhao Luo, Chunhui Zhang +2

Deep trackers have proven success in visual tracking. Typically, these trackers employ optimally pre-trained deep networks to represent all diverse objects with multi-channel featu…

cs.CV2025

Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection

Daichi Zhang, Tong Zhang, Jianmin Bao +2

With the rapid development of generative models, detecting generated fake images to prevent their malicious use has become a critical issue recently. Existing methods frame this ch…

cs.CV2024

Learning Natural Consistency Representation for Face Forgery Video Detection

Daichi Zhang, Zihao Xiao, Shikun Li +3

Face Forgery videos have elicited critical social public concerns and various detectors have been proposed. However, fully-supervised detectors may lead to easily overfitting to sp…

cs.CV2025

COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking

Chunhui Zhang, Li Liu, Jialin Gao +5

Transformer has recently demonstrated great potential in improving vision-language (VL) tracking algorithms. However, most of the existing VL trackers rely on carefully designed me…

cs.CV2024

Low-Resolution Face Recognition via Adaptable Instance-Relation Distillation

Ruixin Shi, Weijia Guo, Shiming Ge

Low-resolution face recognition is a challenging task due to the missing of informative details. Recent approaches based on knowledge distillation have proven that high-resolution…

cs.CV2025

How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline

Chunhui Zhang, Li Liu, Zhipeng Zhang +5

Unmanned Aerial Vehicles (UAVs) offer wide-ranging applications but also pose significant safety and privacy violation risks in areas like airport and infrastructure inspection, sp…

cs.CV2025

Enhancing Frequency Forgery Clues for Diffusion-Generated Image Detection

Daichi Zhang, Tong Zhang, Shiming Ge +1

Diffusion models have achieved remarkable success in image synthesis, but the generated high-quality images raise concerns about potential malicious use. Existing detectors often s…

cs.CV2026

PKI: Prior Knowledge-Infused Neural Network for Few-Shot Class-Incremental Learning

Kexin Baoa, Fanzhao Lin, Zichen Wang +3

Few-shot class-incremental learning (FSCIL) aims to continually adapt a model on a limited number of new-class examples, facing two well-known challenges: catastrophic forgetting a…

cs.CV2024

DANCE: Dual-View Distribution Alignment for Dataset Condensation

Hansong Zhang, Shikun Li, Fanzhao Lin +3

Dataset condensation addresses the problem of data burden by learning a small synthetic training set that preserves essential knowledge from the larger real training set. To date,…

cs.CV2022

WebUAV-3M: A Benchmark for Unveiling the Power of Million-Scale Deep UAV Tracking

Chunhui Zhang, Guanjie Huang, Li Liu +5

Unmanned aerial vehicle (UAV) tracking is of great significance for a wide range of applications, such as delivery and agriculture. Previous benchmarks in this area mainly focused…

cs.LG2024

Towards Personalized Federated Learning via Comprehensive Knowledge Distillation

Pengju Wang, Bochao Liu, Weijia Guo +2

Federated learning is a distributed machine learning paradigm designed to protect data privacy. However, data heterogeneity across various clients results in catastrophic forgettin…

cs.LG2024

Learning Privacy-Preserving Student Networks via Discriminative-Generative Distillation

Shiming Ge, Bochao Liu, Pengju Wang +2

While deep models have proved successful in learning rich knowledge from massive well-annotated data, they may pose a privacy leakage risk in practical deployment. It is necessary…

cs.LG2024

Privacy-Preserving Student Learning with Differentially Private Data-Free Distillation

Bochao Liu, Jianghu Lu, Pengju Wang +4

Deep learning models can achieve high inference accuracy by extracting rich knowledge from massive well-annotated data, but may pose the risk of data privacy leakage in practical d…

cs.LG2024

Learning Differentially Private Diffusion Models via Stochastic Adversarial Distillation

Bochao Liu, Pengju Wang, Shiming Ge

While the success of deep learning relies on large amounts of training datasets, data is often limited in privacy-sensitive domains. To address this challenge, generative model lea…

cs.CV2024

Low-Resolution Object Recognition with Cross-Resolution Relational Contrastive Distillation

Kangkai Zhang, Shiming Ge, Ruixin Shi +1

Recognizing objects in low-resolution images is a challenging task due to the lack of informative details. Recent studies have shown that knowledge distillation approaches can effe…

cs.CV2017

Efficient Privacy Preserving Viola-Jones Type Object Detection via Random Base Image Representation

Xin Jin, Peng Yuan, Xiaodong Li +4

A cloud server spent a lot of time, energy and money to train a Viola-Jones type object detector with high accuracy. Clients can upload their photos to the cloud server to find obj…

cs.CV2026

Few-shot Class-Incremental Learning via Generative Co-Memory Regularization

Kexin Bao, Yong Li, Dan Zeng +1

Few-shot class-incremental learning (FSCIL) aims to incrementally learn models from a small amount of novel data, which requires strong representation and adaptation ability of mod…

cs.CV2017

Predicting Aesthetic Score Distribution through Cumulative Jensen-Shannon Divergence

Xin Jin, Le Wu, Xiaodong Li +6

Aesthetic quality prediction is a challenging task in the computer vision community because of the complex interplay with semantic contents and photographic technologies. Recent st…

cs.CV2019

Low-resolution Face Recognition in the Wild via Selective Knowledge Distillation

Shiming Ge, Shengwei Zhao, Chenyu Li +1

Typically, the deployment of face recognition models in the wild needs to identify low-resolution faces with extremely low computational cost. To address this problem, a feasible s…

cs.CR2023

Model Conversion via Differentially Private Data-Free Distillation

Bochao Liu, Pengju Wang, Shikun Li +2

While massive valuable deep models trained on large-scale data have been released to facilitate the artificial intelligence community, they may encounter attacks in deployment whic…

cs.AI2026

Denoising Implicit Feedback for Cold-start Recommendation

Gaode Chen, Shicheng Wang, Shikun Li +8

Implicit feedback is widely used in recommender systems due to its accessibility and generality, yet it usually presents noisy samples (e.g., clickbait, position bias). Meanwhile,…

cs.HC2024

Transferring Annotator- and Instance-dependent Transition Matrix for Learning from Crowds

Shikun Li, Xiaobo Xia, Jiankang Deng +2

Learning from crowds describes that the annotations of training data are obtained with crowd-sourcing services. Multiple annotators each complete their own small part of the annota…

cs.CV2022

Deepfake Video Detection with Spatiotemporal Dropout Transformer

Daichi Zhang, Fanzhao Lin, Yingying Hua +3

While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each fr…

cs.CV2024

Domain Adaptive Attention Learning for Unsupervised Person Re-Identification

Yangru Huang, Peixi Peng, Yi Jin +3

Person re-identification (Re-ID) across multiple datasets is a challenging task due to two main reasons: the presence of large cross-dataset distinctions and the absence of annotat…

cs.CV2019

Aesthetic Attributes Assessment of Images

Xin Jin, Le Wu, Geng Zhao +6

Image aesthetic quality assessment has been a relatively hot topic during the last decade. Most recently, comments type assessment (aesthetic captions) has been proposed to describ…

cs.CV2021

The 2nd Anti-UAV Workshop & Challenge: Methods and Results

Jian Zhao, Gang Wang, Jianan Li +9

The 2nd Anti-UAV Workshop \& Challenge aims to encourage research in developing novel and accurate methods for multi-scale object tracking. The Anti-UAV dataset used for the Anti-U…

cs.LG2024

Federated Learning with Label-Masking Distillation

Jianghu Lu, Shikun Li, Kexin Bao +3

Federated learning provides a privacy-preserving manner to collaboratively train models on data distributed over multiple local clients via the coordination of a global server. In…

cs.CV2017

Single Reference Image based Scene Relighting via Material Guided Filtering

Xin Jin, Yannan Li, Ningning Liu +4

Image relighting is to change the illumination of an image to a target illumination effect without known the original scene geometry, material information and illumination conditio…

cs.CV2022

Bootstrapping Multi-view Representations for Fake News Detection

Qichao Ying, Xiaoxiao Hu, Yangming Zhou +3

Previous researches on multimedia fake news detection include a series of complex feature extraction and fusion networks to gather useful information from the news. However, how cr…

cs.CV2020

Ultrafast Video Attention Prediction with Coupled Knowledge Distillation

Kui Fu, Peipei Shi, Yafei Song +3

Large convolutional neural network models have recently demonstrated impressive performance on video attention prediction. Conventionally, these models are with intensive computati…

cs.LG2023

Multi-Label Noise Transition Matrix Estimation with Label Correlations: Theory and Algorithm

Shikun Li, Xiaobo Xia, Hansong Zhang +2

Noisy multi-label learning has garnered increasing attention due to the challenges posed by collecting large-scale accurate labels, making noisy labels a more practical alternative…

cs.LG2026

Privacy-Preserving Model Transcription with Differentially Private Synthetic Distillation

Bochao Liu, Shiming Ge, Pengju Wang +2

While many deep learning models trained on private datasets have been deployed in various practical tasks, they may pose a privacy leakage risk as attackers could recover informati…

cs.CV2019

Focusing and Diffusion: Bidirectional Attentive Graph Convolutional Networks for Skeleton-based Action Recognition

Jialin Gao, Tong He, Xi Zhou +1

A collection of approaches based on graph convolutional networks have proven success in skeleton-based action recognition by exploring neighborhood information and dense dependenci…

cs.CV2018

ILGNet: Inception Modules with Connected Local and Global Features for Efficient Image Aesthetic Quality Classification using Domain Adaptation

Xin Jin, Le Wu, Xiaodong Li +6

In this paper, we address a challenging problem of aesthetic image classification, which is to label an input image as high or low aesthetic quality. We take both the local and glo…

cs.CV2024

Look Through Masks: Towards Masked Face Recognition with De-Occlusion Distillation

Chenyu Li, Shiming Ge, Daichi Zhang +1

Many real-world applications today like video surveillance and urban governance need to address the recognition of masked faces, where content replacement by diverse masks often br…

cs.CV2026

CD^2: Constrained Dataset Distillation for Few-Shot Class-Incremental Learning

Kexin Bao, Daichi Zhang, Hansong Zhang +3

Few-shot class-incremental learning (FSCIL) receives significant attention from the public to perform classification continuously with a few training samples, which suffers from th…

cs.CV2026

Divide and Conquer: Static-Dynamic Collaboration for Few-Shot Class-Incremental Learning

Kexin Bao, Daichi Zhang, Yong Li +2

Few-shot class-incremental learning (FSCIL) aims to continuously recognize novel classes under limited data, which suffers from the key stability-plasticity dilemma: balancing the…

cs.CV2024

Interpret the Predictions of Deep Networks via Re-Label Distillation

Yingying Hua, Shiming Ge, Daichi Zhang

Interpreting the predictions of a black-box deep network can facilitate the reliability of its deployment. In this work, we propose a re-label distillation approach to learn a dire…

cs.LG2022

Trustable Co-label Learning from Multiple Noisy Annotators

Shikun Li, Tongliang Liu, Jiyong Tan +2

Supervised deep learning depends on massive accurately annotated examples, which is usually impractical in many real-world scenarios. A typical alternative is learning from multipl…

cs.CV2017

Privacy Preserving Face Retrieval in the Cloud for Mobile Users

Xin Jin, Shiming Ge, Chenggen Song

Recently, cloud storage and processing have been widely adopted. Mobile users in one family or one team may automatically backup their photos to the same shared cloud storage space…

cs.LG2021

Student Network Learning via Evolutionary Knowledge Distillation

Kangkai Zhang, Chunhui Zhang, Shikun Li +2

Knowledge distillation provides an effective way to transfer knowledge via teacher-student learning, where most existing distillation approaches apply a fixed pre-trained model as…

cs.CV2021

Interpretable Face Manipulation Detection via Feature Whitening

Yingying Hua, Daichi Zhang, Pengju Wang +1

Why should we trust the detections of deep neural networks for manipulated faces? Understanding the reasons is important for users in improving the fairness, reliability, privacy a…

cs.CV2020

Receptive Multi-granularity Representation for Person Re-Identification

Guanshuo Wang, Yufeng Yuan, Jiwei Li +2

A key for person re-identification is achieving consistent local details for discriminative representation across variable environments. Current stripe-based feature learning appro…

cs.LG2022

Robust Weight Perturbation for Adversarial Training

Chaojian Yu, Bo Han, Mingming Gong +4

Overfitting widely exists in adversarial robust training of deep networks. An effective remedy is adversarial weight perturbation, which injects the worst-case weight perturbation…

cs.CV2022

Selective-Supervised Contrastive Learning with Noisy Labels

Shikun Li, Xiaobo Xia, Shiming Ge +1

Deep networks have strong capacities of embedding data into latent representations and finishing following tasks. However, the capacities largely come from high-quality annotated l…

cs.LG2024

Personalized Federated Learning via Backbone Self-Distillation

Pengju Wang, Bochao Liu, Dan Zeng +2

In practical scenarios, federated learning frequently necessitates training personalized models for each client using heterogeneous data. This paper proposes a backbone self-distil…

cs.CV2019

Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency

Jia Li, Kui Fu, Shengwei Zhao +1

The performance of video saliency estimation techniques has achieved significant advances along with the rapid development of Convolutional Neural Networks (CNNs). However, devices…

cs.CV2024

Masked Face Recognition with Generative-to-Discriminative Representations

Shiming Ge, Weijia Guo, Chenyu Li +3

Masked face recognition is important for social good but challenged by diverse occlusions that cause insufficient or inaccurate representations. In this work, we propose a unified…

cs.LG2024

Coupled Confusion Correction: Learning from Crowds with Sparse Annotations

Hansong Zhang, Shikun Li, Dan Zeng +2

As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sou…

cs.CV2024

Efficient Low-Resolution Face Recognition via Bridge Distillation

Shiming Ge, Shengwei Zhao, Chenyu Li +2

Face recognition in the wild is now advancing towards light-weight models, fast inference speed and resolution-adapted capability. In this paper, we propose a bridge distillation a…

cs.CV2024

M3D: Dataset Condensation by Minimizing Maximum Mean Discrepancy

Hansong Zhang, Shikun Li, Pengju Wang +2

Training state-of-the-art (SOTA) deep models often requires extensive data, resulting in substantial training and storage costs. To address these challenges, dataset condensation h…

cs.CV2025

Underwater Camouflaged Object Tracking Meets Vision-Language SAM2

Chunhui Zhang, Li Liu, Guanjie Huang +5

Over the past decade, significant progress has been made in visual object tracking, largely due to the availability of large-scale datasets. However, these datasets have primarily…

cs.CV2024

Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition

Junzheng Zhang, Weijia Guo, Bochao Liu +3

Very low-resolution face recognition is challenging due to the serious loss of informative facial details in resolution degradation. In this paper, we propose a generative-discrimi…

cs.CV2024

Private Gradient Estimation is Useful for Generative Modeling

Bochao Liu, Pengju Wang, Weijia Guo +4

While generative models have proved successful in many domains, they may pose a privacy leakage risk in practical deployment. To address this issue, differentially private generati…

cs.CV2020

Accurate Temporal Action Proposal Generation with Relation-Aware Pyramid Network

Jialin Gao, Zhixiang Shi, Jiani Li +4

Accurate temporal action proposals play an important role in detecting actions from untrimmed videos. The existing approaches have difficulties in capturing global contextual infor…

cs.CV2024

Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition

Shiming Ge, Kangkai Zhang, Haolin Liu +4

In spite of great success in many image recognition tasks achieved by recent deep models, directly applying them to recognize low-resolution images may suffer from low accuracy due…