From the 1 of 3 linked papers with an AI index.
3 papers
Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding
Ãnal Ege GaznepoÄlu, Frank Zalkow, Mohammad Joshaghani +3
Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. Vocos is a notable example due to its efficiency, but it…
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes +3
The paper adapts a diffusion-based multi-instrument music synthesis model to perform speech and singing voice conversion by conditioning on phonetic posteriorgrams and pitch contou…
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar +2
In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically…