3 papers
eess.AS2026
Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding
Ünal Ege Gaznepoğlu, Frank Zalkow, Mohammad Joshaghani +3
Recently, time-frequency neural vocoders have been approaching the state-of-the-art quality of time-domain neural vocoders. Vocos is a notable example due to its efficiency, but it…
cs.SD2026
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes +3
Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However,…
eess.AS2025
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar +2
In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically…