collaborators

7 papers

cs.SD2026

QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection

Duc-Tuan Truong, Tianchi Liu, Ruijie Tao +3

Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-ce…

cs.SD2026

Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound

Liumeng Xue, Ziya Zhou, Jiahao Pan +20

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and gene…

eess.AS2026

Evaluating the Expressive Appropriateness of Speech in Rich Contexts

Tianrui Wang, Ziyang Ma, Yizhou Peng +26

Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its con…

cs.CR2026

Scores Know Bobs Voice: Speaker Impersonation Attack

Chanwoo Hwang, Sunpill Kim, Yong Kiam Tan +6

Advances in deep learning have enabled the widespread deployment of speaker recognition systems (SRSs), yet they remain vulnerable to score-based impersonation attacks. Existing at…

cs.SD2026

Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection

Duc-Tuan Truong, Tianchi Liu, Junjie Li +3

In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during tr…

cs.SD2026

Xi+: Uncertainty Supervision for Robust Speaker Embedding

Junjie Li, Kong Aik Lee, Duc-Tuan Truong +2

There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Sinc…