collaborators

5 papers

eess.AS2023

On Feature Importance and Interpretability of Speaker Representations

Frederik Rautenberg, Michael Kuhlmann, Jana Wiechmann +3

Unsupervised speech disentanglement aims at separating fast varying from slowly varying components of a speech signal. In this contribution, we take a closer look at the embedding…

eess.AS2023

Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization

Thilo von Neumann, Christoph Boeddeker, Tobias Cord-Landwehr +2

We propose a modular pipeline for the single-channel separation, recognition, and diarization of meeting-style recordings and evaluate it on the Libri-CSS dataset. Using a Continuo…

cs.SD2023

LibriWASN: A Data Set for Meeting Separation, Diarization, and Recognition with Asynchronous Recording Devices

Joerg Schmalenstroeer, Tobias Gburrek, Reinhold Haeb-Umbach

We present LibriWASN, a data set whose design follows closely the LibriCSS meeting recognition data set, with the marked difference that the data is recorded with devices that are…

eess.AS2023

Investigating Speaker Embedding Disentanglement on Natural Read Speech

Michael Kuhlmann, Adrian Meise, Fritz Seebauer +2

Disentanglement is the task of learning representations that identify and separate factors that explain the variation observed in data. Disentangled representations are useful to i…

cs.CL2023

MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems

Thilo von Neumann, Christoph Boeddeker, Marc Delcroix +1

MeetEval is an open-source toolkit to evaluate all kinds of meeting transcription systems. It provides a unified interface for the computation of commonly used Word Error Rates (WE…