10 papers
Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System
Thomas Thebaud, Sonal Joshi, Henry Li +4
Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack scenario on a speech commands classificatio…
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
Kaavya Chaparala, Thomas Thebaud, Jesús Villalba López +3
There are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias int…
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
Junhyeok Lee, Helin Wang, Yaohan Guan +4
We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models r…
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
Yen-Ju Lu, Kunxiao Gao, Mingrui Liang +5
Recent audio language models can follow long conversations. However, research on emotion-aware or spoken dialogue summarization is constrained by the lack of data that links speech…
Multi-Target Backdoor Attacks Against Speaker Recognition
Alexandrine Fortier, Sonal Joshi, Thomas Thebaud +3
In this work, we propose a multi-target backdoor attack against speaker identification using position-independent clicking sounds as triggers. Unlike previous single-target approac…
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
Xiluo He, Alexander Polok, Jesús Villalba +2
An increasingly common training paradigm for multi-talker automatic speech recognition (ASR) is to use speaker activity signals to adapt single-speaker ASR models for overlapping s…