4 papers · 1 filter
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
Paul Primus, Florian Schmid, Gerhard Widmer
Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text…
Effective Pre-Training of Audio Transformers for Sound Event Detection
Florian Schmid, Tobias Morocutti, Francesco Foscarin +3
We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously design…
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
Paul Primus, Florian Schmid, Gerhard Widmer
Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embeddi…
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
Jonathan Greif, Florian Schmid, Paul Primus +1
Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sou…