Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
arXiv:2309.08454 · doi:10.1109/SLT61566.2024.10832307
Abstract
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free streams on which ASR is performed. Recently, TF-GridNet has shown impressive performance in speech separation in real reverberant conditions. Furthermore, a mixture encoder was proposed that leverages the mixed speech to mitigate the effect of separation artifacts. In this work, we extended the mixture encoder from a static two-speaker scenario to a natural meeting context featuring an arbitrary number of speakers and varying degrees of overlap. We further demonstrate its limits by the integration with separators of varying strength including TF-GridNet. Our experiments result in a new state-of-the-art performance on LibriCSS using a single microphone. They show that TF-GridNet largely closes the gap between previous methods and oracle separation independent of mixture encoding. We further investigate the remaining potential for improvement.
Presented at SLT 2024
References in corpus (11)
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Language Modeling with Deep Transformers
- RWTH ASR Systems for LibriSpeech: Hybrid vs Attention -- w/o Data Augmentation
- Recognizing Overlapped Speech in Meetings: A Multichannel Separation Approach Using Neural Networks
- RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition
- SMS-WSJ: Database, performance measures, and baseline recipe for multi-channel source separation and recognition
- Progressive Joint Modeling in Unsupervised Single-channel Overlapped Speech Recognition
- TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
- On Word Error Rate Definitions and their Efficient Computation for Multi-Speaker Speech Recognition Systems