22 citations · 27 across the 4 of their papers we have counts for
4 papers
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
Parthasaarathy Sudarsanam, Irene Martín-Morató, Tuomas Virtanen
This paper proposes a single-stage training approach that semantically aligns three modalities - audio, visual, and text using a contrastive learning framework. Contrastive trainin…
DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
Samuele Cornell, Janek Ebbers, Constance Douwes +4
The Detection and Classification of Acoustic Scenes and Events Challenge Task 4 aims to advance sound event detection (SED) systems in domestic environments by leveraging training…
Training sound event detection with soft labels from crowdsourced annotations
Irene Martín-Morató, Manu Harju, Paul Ahokas +1
In this paper, we study the use of soft labels to train a system for sound event detection (SED). Soft labels can result from annotations which account for human uncertainty about…
Low-complexity acoustic scene classification in DCASE 2022 Challenge
Irene Martín-Morató, Francesco Paissan, Alberto Ancilotto +5
This paper presents an analysis of the Low-Complexity Acoustic Scene Classification task in DCASE 2022 Challenge. The task was a continuation from the previous years, but the low-c…