collaborators

8 papers

cs.SD2026

Time-Frequency Consistency Learning for Robust Speech Deepfake Detection

Jun Xue, Zhuolin Yi, Yanzhen Ren +6

Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lac…

cs.SD2026

Exploring the Scale and Diversity of Speech Anti-spoofing Datasets: Experiments and Analysis

Zhuolin Yi, Jun Xue, Yanzhen Ren +5

The scale of speech anti-spoofing datasets has grown exponentially over the past decade, driven by the assumption that larger data leads to better performance. However, it remains…

cs.SD2026

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

Jun Xue, Yi Chai, Yanzhen Ren +6

Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting in restricted diversity and poor cover…

cs.CV2026

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

Jiayu Xiong, Jing Wang, Qi Zhang +2

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forg…

cs.SD2026

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection

Jun Xue, Tong Zhang, Zhuolin Yi +4

The rapid advancement of generative AI has made audio deepfakes increasingly indistinguishable from authentic human vocals, posing significant threats to persons-of-interest (POI)…

cs.CV2026

When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse

Yihuan Huang, Jun Xue, Liu Jiajun +5

Audio-Visual Speech Recognition (AVSR) has achieved remarkable progress in offline conditions, yet its robustness in real-world video conferencing (VC) remains largely unexplored.…