6 papers · 1 filter
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
Michael Kuhlmann, Alexander Werning, Thilo von Neumann +1
A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cann…
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4
Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…
Error Analysis in a Modular Meeting Transcription System
Peter Vieting, Simon Berger, Thilo von Neumann +3
Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previousl…
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
Thilo von Neumann, Christoph Boeddeker, Marc Delcroix +1
The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker spe…
Utterance-by-utterance overlap-aware neural diarization with Graph-PIT
Keisuke Kinoshita, Thilo von Neumann, Marc Delcroix +2
Recent speaker diarization studies showed that integration of end-to-end neural diarization (EEND) and clustering-based diarization is a promising approach for achieving state-of-t…
All-neural online source separation, counting, and diarization for meeting analysis
Thilo von Neumann, Keisuke Kinoshita, Marc Delcroix +3
Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all…