6 papers
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
Peter Vieting, Simon Berger, Thilo von Neumann +3
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free st…
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
Tobias Cord-Landwehr, Christoph Boeddeker, Reinhold Haeb-Umbach
We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separ…
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
Reinhold Haeb-Umbach, Tomohiro Nakatani, Marc Delcroix +2
Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal…
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
Samuele Cornell, Taejin Park, Steve Huang +6
This paper presents the CHiME-8 DASR challenge which carries on from the previous edition CHiME-7 DASR (C7DASR) and the past CHiME-6 challenge. It focuses on joint multi-channel di…
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
Christoph Boeddeker, Tobias Cord-Landwehr, Reinhold Haeb-Umbach
Diarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker. Particular…
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
Thilo von Neumann, Christoph Boeddeker, Tobias Cord-Landwehr +2
We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuo…