5 papers · 1 filter
GenVC: Self-Supervised Zero-Shot Voice Conversion
Zexin Cai, Henry Li Xinyuan, Ashi Garg +5
Most current zero-shot voice conversion methods rely on externally supervised components, particularly speaker encoders, for training. To explore alternatives that eliminate this d…
Scalable Controllable Accented TTS
Henry Li Xinyuan, Zexin Cai, Ashi Garg +5
We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for…
ShiftySpeech: A Large-Scale Synthetic Speech Dataset with Distribution Shifts
Ashi Garg, Zexin Cai, Lin Zhang +6
The problem of synthetic speech detection has enjoyed considerable attention, with recent methods achieving low error rates across several established benchmarks. However, to what…
HLTCOE JHU Submission to the Voice Privacy Challenge 2024
Henry Li Xinyuan, Zexin Cai, Ashi Garg +5
We present a number of systems for the Voice Privacy Challenge, including voice conversion based systems such as the kNN-VC method and the WavLM voice Conversion method, and text-t…
Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
Zexin Cai, Henry Li Xinyuan, Ashi Garg +5
Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has…