activity
20242026
collaborators

7 papers

cs.SD2026

AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection

Yuankun Xie, Haonan Cheng, Jiayi Zhou +11

Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake…

cs.SD2026

AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

Yuankun Xie, Haonan Cheng, Jiayi Zhou +11

This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realis…

cs.SD2026

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan

Yuankun Xie, Haonan Cheng, Jiayi Zhou +10

The rapid advancement of Audio Large Language Models (ALLMs) has enabled cost-effective, high-fidelity generation and manipulation of both speech and non-speech audio, including so…

cs.CV2025

ResAD++: Towards Class Agnostic Anomaly Detection via Residual Feature Learning

Xincheng Yao, Chao Shi, Muming Zhao +2

This paper explores the problem of class-agnostic anomaly detection (AD), where the objective is to train one class-agnostic AD model that can generalize to detect anomalies in div…

cs.CV2025

Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion

Zesheng Wang, Alexandre Bruckert, Patrick Le Callet +1

Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing app…

cs.MM2025

Perceptual Visual Quality Assessment: Principles, Methods, and Future Directions

Wei Zhou, Hadi Amirpour, Christian Timmerer +3

As multimedia services such as video streaming, video conferencing, virtual reality (VR), and online gaming continue to expand, ensuring high perceptual visual quality becomes a pr…