cross-domain audio generation 1diffusion models 1singing voice conversion 1speech synthesis 1voice conversion 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.SD2026
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Ben Maman, Frank Zalkow, Hans-Ulrich Berendes +3
The paper adapts a diffusion-based multi-instrument music synthesis model to perform speech and singing voice conversion by conditioning on phonetic posteriorgrams and pitch contou…
cs.SD2026
Snapping Matters: Context-Aware Onset Refinement for Automatic Music Transcription
Abhirup Saha, Hans-Ulrich Berendes, Meinard Müller +1
Precise note-level annotations are critical for training automatic music transcription (AMT) systems, in particular note-onset labels, which form a core component of many recent AM…
cs.SD2025
Count The Notes: Histogram-Based Supervision for Automatic Music Transcription
Jonathan Yaffe, Ben Maman, Meinard Müller +1
Automatic Music Transcription (AMT) converts audio recordings into symbolic musical representations. Training deep neural networks (DNNs) for AMT typically requires strongly aligne…