3 papers
eess.AS2024
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partiall…
eess.AS2023
Frame-wise and overlap-robust speaker embeddings for meeting diarization
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact…
eess.AS2023
Self-regularised Minimum Latency Training for Streaming Transformer-based Speech Recognition
Mohan Li, Rama Doddipatla, Catalin Zorila
This paper proposes a self-regularised minimum latency training (SR-MLT) method for streaming Transformer-based automatic speech recognition (ASR) systems. In previous works, laten…