5 papers
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Mingrui Liang, Thomas Thebaud, Lukasz Wojciak +4
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing…
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Thomas Thebaud, Junhyeok Lee, Laureano Moro-Velazquez +2
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather t…
StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
Yuzhe Wang, Thomas Thebaud, Jennifer Hu +5
Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited. We introduce StanceB…
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
Kaavya Chaparala, Thomas Thebaud, Jesús Villalba López +3
There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias int…
Multi-Target Backdoor Attacks Against Speaker Recognition
Alexandrine Fortier, Sonal Joshi, Thomas Thebaud +3
In this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approac…