most citedSMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition

56 citations · 59 across the 3 of their papers we have counts for

collaborators

7 papers

eess.AS2020

Far-Field Automatic Speech Recognition

Reinhold Haeb-Umbach, Jahn Heymann, Lukas Drude +3

The machine recognition of speech spoken at a distance from the microphones, known as far-field automatic speech recognition (ASR), has received a significant increase of attention…

eess.AS2020

Multi-talker ASR for an unknown number of sources: Joint training of source counting, separation and ASR

Thilo von Neumann, Christoph Boeddeker, Lukas Drude +4

Most approaches to multi-talker overlapped speech separation and recognition assume that the number of simultaneously active speakers is given, but in realistic situations, it is t…

eess.AS2019

End-to-end training of time domain audio separation and recognition

Thilo von Neumann, Keisuke Kinoshita, Lukas Drude +4

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now,…

cs.SD2019

Demystifying TasNet: A Dissecting Approach

Jens Heitkaemper, Darius Jakobeit, Christoph Boeddeker +2

In recent years time domain speech separation has excelled over frequency domain separation in single channel scenarios and noise-free environments. In this paper we dissect the ga…

cs.SD201956 cited

SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition

Lukas Drude, Jens Heitkaemper, Christoph Boeddeker +1

We present a multi-channel database of overlapping speech for training, evaluation, and detailed analysis of source separation and extraction algorithms: SMS-WSJ -- Spatialized Mul…

cs.SD2019

Unsupervised training of neural mask-based beamforming

Lukas Drude, Jahn Heymann, Reinhold Haeb-Umbach

We present an unsupervised training approach for a neural network-based mask estimator in an acoustic beamforming application. The network is trained to maximize a likelihood crite…