activity
20122021
most citedAddressing Missing Labels in Large-Scale Sound Event Recognition Using a Teacher-Student Framework With Loss Masking

23 citations · 42 across the 4 of their papers we have counts for

collaborators

7 papers

cs.SD2021

The Benefit Of Temporally-Strong Labels In Audio Event Classification

Shawn Hershey, Daniel P W Ellis, Eduardo Fonseca +4

To reveal the importance of temporal precision in ground truth audio event labels, we collected precise (~0.1 sec resolution) "strong" labels for a portion of the AudioSet dataset.…

cs.SD2021

Self-Supervised Learning from Automatically Separated Sound Scenes

Eduardo Fonseca, Aren Jansen, Daniel P. W. Ellis +7

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The associati…

cs.SD2020

Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

Efthymios Tzinis, Scott Wisdom, Aren Jansen +4

Recent progress in deep learning has enabled many advances in sound separation and visual scene understanding. However, extracting sound sources which are apparent in natural video…

cs.SD202023 cited

Addressing Missing Labels in Large-Scale Sound Event Recognition Using a Teacher-Student Framework With Loss Masking

Eduardo Fonseca, Shawn Hershey, Manoj Plakal +4

The study of label noise in sound event recognition has recently gained attention with the advent of larger and noisier datasets. This work addresses the problem of missing labels,…

cs.SD2019

Coincidence, Categorization, and Consolidation: Learning to Recognize Sounds with Minimal Supervision

Aren Jansen, Daniel P. W. Ellis, Shawn Hershey +4

Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-lab…

cs.SD20175 cited

Unsupervised Learning of Semantic Audio Representations

Aren Jansen, Manoj Plakal, Ratheet Pandya +5

Even in the absence of any explicit semantic annotation, vast collections of audio recordings provide valuable information for learning the categorical structure of sounds. We cons…