4 papers
Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
Alexander Fichtinger, Jan Schlüter, Gerhard Widmer
Generative models of music audio are typically used to generate output based solely on a text prompt or melody. Boomerang sampling, recently proposed for the image domain, allows g…
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
Paul Primus, Florian Schmid, Gerhard Widmer
Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text…
Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
Tobias Morocutti, Florian Schmid, Khaled Koutini +1
Knowledge Distillation (KD) is a widespread technique for compressing the knowledge of large models into more compact and efficient models. KD has proved to be highly effective in…
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
Tobias Morocutti, Florian Schmid, Jonathan Greif +2
We target the problem of developing new low-complexity networks for the sound event detection task. Our goal is to meticulously analyze the performance-complexity trade-off, aiming…