1 citations · 1 across the 5 of their papers we have counts for
5 papers
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
Paul Primus, Florian Schmid, Gerhard Widmer
Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embeddi…
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
Jonathan Greif, Florian Schmid, Paul Primus +1
Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sou…
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
Florian Schmid, Paul Primus, Tobias Morocutti +2
This technical report describes the CP-JKU team's submission for Task 4 Sound Event Detection with Heterogeneous Training Datasets and Potentially Missing Labels of the DCASE 24 Ch…
Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
Florian Schmid, Paul Primus, Tobias Morocutti +2
A central problem in building effective sound event detection systems is the lack of high-quality, strongly annotated sound event datasets. For this reason, Task 4 of the DCASE 202…
Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models
Florian Schmid, Khaled Koutini, Gerhard Widmer
The introduction of large-scale audio datasets, such as AudioSet, paved the way for Transformers to conquer the audio domain and replace CNNs as the state-of-the-art neural network…