31 citations · 71 across the 5 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023★ 1 cited
ACE-VC: Adaptive and Controllable Voice Conversion using Explicitly Disentangled Self-supervised Speech Representations
Shehzeen Hussain, Paarth Neekhara, Jocelyn Huang +2
In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a…
cs.SD2019
Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens
Rafael Valle, Jason Li, Ryan Prenger +1
Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data. By explicitly conditioning…