28 papers
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Mingrui Liang, Thomas Thebaud, Lukasz Wojciak +4
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing…
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez +2
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather t…
StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
Yuzhe Wang, Thomas Thebaud, Jennifer Hu +5
Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceB…
Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System
Thomas Thebaud, Sonal Joshi, Henry Li +4
Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a speech commands classificatio…
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
Thomas Thebaud, Yuzhe Wang, Laureano Moro-Velazquez +2
Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or specific fields such as emotions or the sp…
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
Alexandrine Fortier, Thomas Thebaud, Jesús Villalba +3
Speech language models (SLMs) are systems of systems: independent components that unite to achieve a common goal. Despite their heterogeneous nature, SLMs are often studied end-to-…