7 papers
QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
Duc-Tuan Truong, Tianchi Liu, Ruijie Tao +3
Recent work shows that one-class learning can detect unseen deepfake attacks by modeling a compact distribution of bona fide speech around a single centroid. However, the single-ce…
Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound
Liumeng Xue, Ziya Zhou, Jiahao Pan +20
Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and gene…
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang, Ziyang Ma, Yizhou Peng +26
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its con…
Scores Know Bobs Voice: Speaker Impersonation Attack
Chanwoo Hwang, Sunpill Kim, Yong Kiam Tan +6
Advances in deep learning have enabled the widespread deployment of speaker recognition systems (SRSs), yet they remain vulnerable to score-based impersonation attacks. Existing at…
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
Duc-Tuan Truong, Tianchi Liu, Junjie Li +3
In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during tr…
Xi+: Uncertainty Supervision for Robust Speaker Embedding
Junjie Li, Kong Aik Lee, Duc-Tuan Truong +2
There are various factors that can influence the performance of speaker recognition systems, such as emotion, language and other speaker-related or context-related variations. Sinc…