most citedWord Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition

2 citations · 2 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2026

Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization

Marc Deegen, Adrian Meise, Reinhold Haeb-Umbach

Spatial information gleaned from multi-channel input has been shown to lead to improvements in meeting processing tasks like diarization and source separation. At the same time, di…

eess.AS2026

Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments

Adrian Meise, Tobias Cord-Landwehr, Christoph Boeddeker +3

Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in mee…

eess.AS2026

On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization

Marc Deegen, Tobias Gburrek, Tobias Cord-Landwehr +4

Recent advances in speaker diarization exploit large pretrained foundation models, such as WavLM, to achieve state-of-the-art performance on multiple datasets. Systems like DiariZe…

eess.AS2025

Error Analysis in a Modular Meeting Transcription System

Peter Vieting, Simon Berger, Thilo von Neumann +3

Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previousl…

eess.AS20252 cited

Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition

Thilo von Neumann, Christoph Boeddeker, Marc Delcroix +1

The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker spe…