activity
20162022
most citedMultichannel End-to-end Speech Recognition

46 citations · 102 across the 11 of their papers we have counts for

collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS2022

CycleGAN-Based Unpaired Speech Dereverberation

Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan +4

Typically, neural network-based speech dereverberation models are trained on paired data, composed of a dry utterance and its corresponding reverberant utterance. The main limitati…

eess.AS20213 cited

Improving Bird Classification with Unsupervised Sound Separation

Tom Denton, Scott Wisdom, John R. Hershey

This paper addresses the problem of species classification in bird song recordings. The massive amount of available field recordings of birds presents an opportunity to use machine…

eess.AS20217 cited

DF-Conformer: Integrated architecture of Conv-TasNet and Conformer using linear complexity self-attention for speech enhancement

Yuma Koizumi, Shigeki Karita, Scott Wisdom +4

Single-channel speech enhancement (SE) is an important task in speech processing. A widely used framework combines an analysis/synthesis filterbank with a mask prediction network,…

eess.AS2021

Sparse, Efficient, and Semantic Mixture Invariant Training: Taming In-the-Wild Unsupervised Sound Separation

Scott Wisdom, Aren Jansen, Ron J. Weiss +2

Supervised neural network training has led to significant progress on single-channel sound separation. This approach relies on ground truth isolated sources, which precludes scalin…

eess.AS20206 cited

Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording

Cong Han, Yi Luo, Chenda Li +8

Leveraging additional speaker information to facilitate speech separation has received increasing attention in recent years. Recent research includes extracting target speech by us…

eess.AS2020

Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis

Desh Raj, Pavel Denisov, Zhuo Chen +11

Multi-speaker speech recognition of unsegmented recordings has diverse applications such as meeting transcription and automatic subtitle generation. With technical advances in syst…