5 papers
On Feature Importance and Interpretability of Speaker Representations
Frederik Rautenberg, Michael Kuhlmann, Jana Wiechmann +3
Unsupervised speech disentanglement aims at separating fast varying from slowly varying components of a speech signal. In this contribution, we take a closer look at the embedding…
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
Thilo von Neumann, Christoph Boeddeker, Tobias Cord-Landwehr +2
We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuo…
LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices
Joerg Schmalenstroeer, Tobias Gburrek, Reinhold Haeb-Umbach
We present LibriWASN, a data set whose design follows closely the LibriCSS meeting recognition data set, with the marked difference that the data is recorded with devices that are…
Investigating Speaker Embedding Disentanglement on Natural Read Speech
Michael Kuhlmann, Adrian Meise, Fritz Seebauer +2
Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to i…
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
Thilo von Neumann, Christoph Boeddeker, Marc Delcroix +1
MeetEval is an open-source toolkit to evaluate all kinds of meeting transcription systems. It provides a unified interface for the computation of commonly used Word Error Rates (WE…