6 papers
Multimodal Speaker Verification as a Threat to Speaker Anonymization
Ashi Garg, Cristina Aggazzotti, Leibny Paola GarcÃa-Perera +1
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulat…
Content Anonymization for Privacy in Long-form Audio
Cristina Aggazzotti, Ashi Garg, Zexin Cai +1
Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmarks such as the VoicePrivacy Challenge.…
GenVC: Self-Supervised Zero-Shot Voice Conversion
Zexin Cai, Henry Li Xinyuan, Ashi Garg +5
Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this d…
Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
Ashi Garg, Zexin Cai, Henry Li Xinyuan +5
We address the challenge of detecting synthesized speech under distribution shifts -- arising from unseen synthesis methods, speakers, languages, or audio conditions -- relative to…
Scalable Controllable Accented TTS
Henry Li Xinyuan, Zexin Cai, Ashi Garg +5
We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for…
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
Ashi Garg, Zexin Cai, Lin Zhang +6
The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what…