cross-domain audio generation 1diffusion models 1singing voice conversion 1speech synthesis 1voice conversion 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.SD2026
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes +3
The paper adapts a diffusion-based multi-instrument music synthesis model to perform speech and singing voice conversion by conditioning on phonetic posteriorgrams and pitch contou…
eess.AS2025
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron
Kishor Kayyar Lakshminarayana, Frank Zalkow, Christian Dittmar +2
In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically…