activity
20182025
most citedHEAR: Holistic Evaluation of Audio Representations

36 citations · 114 across the 10 of their papers we have counts for

collaborators

21 papers

cs.SD2025

Recomposer: Event-roll-guided generative audio editing

Daniel P. W. Ellis, Eduardo Fonseca, Ron J. Weiss +7

Editing complex real-world sound scenes is difficult because individual sound sources overlap in time. Generative models can fill-in missing or corrupted details based on their str…

cs.LG2023★ 11 cited

Dataset balancing can hurt model performance

R. Channing Moore, Daniel P. W. Ellis, Eduardo Fonseca +3

Machine learning from training data with a skewed distribution of examples per class can lead to models that favor performance on common classes at the expense of performance on ra…

cs.CV2022

Audiovisual Masked Autoencoders

Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu +3

Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining archi…

eess.AS2022★ 3 cited

Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system

Francesca Ronchini, Samuele Cornell, Romain Serizel +3

The aim of the Detection and Classification of Acoustic Scenes and Events Challenge Task 4 is to evaluate systems for the detection of sound events in domestic environments using a…

cs.SD2022★ 36 cited

HEAR: Holistic Evaluation of Audio Representations

Joseph Turian, Jordie Shier, Humair Raj Khan +20

What audio embedding approach generalizes best to a wide range of downstream tasks across a variety of everyday domains without fine-tuning? The aim of the HEAR benchmark is to dev…

cs.SD2021★ 4 cited

Improving Sound Event Classification by Increasing Shift Invariance in Convolutional Neural Networks

Eduardo Fonseca, Andres Ferraro, Xavier Serra

Recent studies have put into question the commonly assumed shift invariance property of convolutional networks, showing that small shifts in the input can affect the output predict…