3 papers
eess.AS2025
Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
Patricia Hu, Silvan David Peter, Jan Schlüter +1
Advances in neural network design and the availability of large-scale labeled datasets have driven major improvements in piano transcription. Existing approaches target either offl…
cs.SD2025
Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
Alexander Fichtinger, Jan Schlüter, Gerhard Widmer
Generative models of music audio are typically used to generate output based solely on a text prompt or melody. Boomerang sampling, recently proposed for the image domain, allows g…
eess.AS2024
Effective Pre-Training of Audio Transformers for Sound Event Detection
Florian Schmid, Tobias Morocutti, Francesco Foscarin +3
We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously design…