From the 1 of 11 linked papers with an AI index.
10 papers · 1 filter
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Florian Schmid, Paul Primus, Alexander Fichtinger +3
This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each reco…
Low-Complexity Acoustic Scene Classification with Device Information in the DCASE 2025 Challenge
Florian Schmid, Paul Primus, Toni Heittola +3
This paper presents the Low-Complexity Acoustic Scene Classification with Device Information Task of the DCASE 2025 Challenge, along with its baseline system. Continuing the focus…
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
Paul Primus, Florian Schmid, Gerhard Widmer
Learning to associate audio with textual descriptions is valuable for a range of tasks, including pretraining, zero-shot classification, audio retrieval, audio captioning, and text…
Effective Pre-Training of Audio Transformers for Sound Event Detection
Florian Schmid, Tobias Morocutti, Francesco Foscarin +3
We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously design…
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
Paul Primus, Florian Schmid, Gerhard Widmer
Dual-encoder-based audio retrieval systems are commonly optimized with contrastive learning on a set of matching and mismatching audio-caption pairs. This leads to a shared embeddi…
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
Jonathan Greif, Florian Schmid, Paul Primus +1
Query-by-Vocal Imitation (QBV) is about searching audio files within databases using vocal imitations created by the user's voice. Since most humans can effectively communicate sou…