20 citations · 33 across the 9 of their papers we have counts for
10 papers · 1 filter
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
Michael Kuhlmann, Alexander Werning, Thilo von Neumann +1
A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cann…
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4
Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…
Error Analysis in a Modular Meeting Transcription System
Peter Vieting, Simon Berger, Thilo von Neumann +3
Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previousl…
MMS-MSG: A Multi-purpose Multi-Speaker Mixture Signal Generator
Tobias Cord-Landwehr, Thilo von Neumann, Christoph Boeddeker +1
The scope of speech enhancement has changed from a monolithic view of single, independent tasks, to a joint processing of complex conversational speech recordings. Training and eva…
A Meeting Transcription System for an Ad-Hoc Acoustic Sensor Network
Tobias Gburrek, Christoph Boeddeker, Thilo von Neumann +3
We propose a system that transcribes the conversation of a typical meeting scenario that is captured by a set of initially unsynchronized microphone arrays at unknown positions. It…
Speeding Up Permutation Invariant Training for Source Separation
Thilo von Neumann, Christoph Boeddeker, Keisuke Kinoshita +2
Permutation invariant training (PIT) is a widely used training criterion for neural network-based source separation, used for both utterance-level separation with utterance-level P…