4 papers
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
Tobias Cord-Landwehr, Christoph Boeddeker, Reinhold Haeb-Umbach
We propose an approach for simultaneous diarization and separation of meeting data. It consists of a complex Angular Central Gaussian Mixture Model (cACGMM) for speech source separ…
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
Yuying Xie, Michael Kuhlmann, Frederik Rautenberg +2
Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important…
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
Christoph Boeddeker, Tobias Cord-Landwehr, Reinhold Haeb-Umbach
Diarization is a crucial component in meeting transcription systems to ease the challenges of speech enhancement and attribute the transcriptions to the correct speaker. Particular…
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
We propose a modified teacher-student training for the extraction of frame-wise speaker embeddings that allows for an effective diarization of meeting scenarios containing partiall…