TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
arXiv:2303.03849 · doi:10.1109/TASLP.2024.3350887
Abstract
Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) diarization approach, which assumes that initial speaker embeddings are available. We replace the final combined speaker activity estimation network of TS-VAD with a network that produces speaker activity estimates at a time-frequency resolution. Those act as masks for source extraction, either via masking or via beamforming. The technique can be applied both for single-channel and multi-channel input and, in both cases, achieves a new state-of-the-art word error rate (WER) on the LibriCSS meeting data recognition task. We further compute speaker-aware and speaker-agnostic WERs to isolate the contribution of diarization errors to the overall WER performance.
Submitted to IEEE/ACM TASLP
References in corpus (6)
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario
- Auto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap
- On Word Error Rate Definitions and their Efficient Computation for Multi-Speaker Speech Recognition Systems
Cited by in corpus (4)
- PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
- Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
- Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
- Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder