most citedDynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models

1 citations · 1 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval

Paul Primus, Florian Schmid, Gerhard Widmer

Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embeddi…

eess.AS2024

Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining

Jonathan Greif, Florian Schmid, Paul Primus +1

Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sou…

eess.AS2024

Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training

Florian Schmid, Paul Primus, Tobias Morocutti +2

This technical report describes the CP-JKU team's submission for Task 4 Sound Event Detection with Heterogeneous Training Datasets and Potentially Missing Labels of the DCASE 24 Ch…

eess.AS2024

Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets

Florian Schmid, Paul Primus, Tobias Morocutti +2

A central problem in building effective sound event detection systems is the lack of high-quality, strongly annotated sound event datasets. For this reason, Task 4 of the DCASE 202…

cs.SD20231 cited

Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models

Florian Schmid, Khaled Koutini, Gerhard Widmer

The introduction of large-scale audio datasets, such as AudioSet, paved the way for Transformers to conquer the audio domain and replace CNNs as the state-of-the-art neural network…