cross-domain audio generation 1diffusion models 1singing voice conversion 1speech synthesis 1voice conversion 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.SD2026
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes +3
The paper adapts a diffusion-based multi-instrument music synthesis model to perform speech and singing voice conversion by conditioning on phonetic posteriorgrams and pitch contou…
cs.SD2026
Snapping Matters: Context-Aware Onset Refinement for Automatic Music Transcription
Abhirup Saha, Hans-Ulrich Berendes, Meinard Müller +1
Precise note-level annotations are critical for training automatic music transcription (AMT) systems, in particular note-onset labels, which form a core component of many recent AM…