4 papers
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
Michael Kuhlmann, Alexander Werning, Thilo von Neumann +1
A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cann…
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4
Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…
Error Analysis in a Modular Meeting Transcription System
Peter Vieting, Simon Berger, Thilo von Neumann +3
Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previousl…
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
Peter Vieting, Simon Berger, Thilo von Neumann +3
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free st…