activity
20172025
most citedSound Event Detection and Separation: a Benchmark on Desed Synthetic Soundscapes

30 citations · 65 across the 16 of their papers we have counts for

collaborators
Showing 2021Show all

7 papers · 1 filter

cs.SD2021

Adapting Speech Separation to Real-World Meetings Using Mixture Invariant Training

Aswin Sivaraman, Scott Wisdom, Hakan Erdogan +1

The recently-proposed mixture invariant training (MixIT) is an unsupervised method for training single-channel sound separation models in the sense that it does not require ground-…

eess.AS20213 cited

Improving Bird Classification with Unsupervised Sound Separation

Tom Denton, Scott Wisdom, John R. Hershey

This paper addresses the problem of species classification in bird song recordings. The massive amount of available field recordings of birds presents an opportunity to use machine…

eess.AS20217 cited

DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement

Yuma Koizumi, Shigeki Karita, Scott Wisdom +4

Single-channel speech enhancement (SE) is an important task in speech processing. A widely used framework combines an analysis/synthesis filterbank with a mask prediction network,…

cs.SD20213 cited

Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention

Efthymios Tzinis, Scott Wisdom, Tal Remez +1

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking…

eess.AS2021

Sparse, Efficient, and Semantic Mixture Invariant Training: Taming In-the-Wild Unsupervised Sound Separation

Scott Wisdom, Aren Jansen, Ron J. Weiss +2

Supervised neural network training has led to significant progress on single-channel sound separation. This approach relies on ground truth isolated sources, which precludes scalin…

cs.SD2021

End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings

Soumi Maiti, Hakan Erdogan, Kevin Wilson +3

We present an end-to-end deep network model that performs meeting diarization from single-channel audio recordings. End-to-end diarization models have the advantage of handling spe…