81 citations · 108 across the 9 of their papers we have counts for
1 paper · 2 filters
Rafael Valle, Jason Li, Ryan Prenger +1
Mellotron is a multispeaker voice synthesis model based on Tacotron 2 GST that can make a voice emote and sing without emotive or singing training data. By explicitly conditioning…