From the 1 of 507 papers with an AI index.
29.3k citations
- University of ArizonaUS20 papers
- Complexity and Topology in Quantum MatterDE15 papers
- Max Planck Institute for the Science of LightDE14 papers
- Ruhr University BochumDE14 papers
- TU Dortmund UniversityDE13 papers
- Beijing Institute of TechnologyCN11 papers
- Centre National de la Recherche ScientifiqueFR11 papers
- Leipzig UniversityDE10 papers
- Technical University of MunichDE10 papers
- University of WürzburgDE10 papers
- Forschungszentrum JülichDE8 papers
- Technische Universität DresdenDE8 papers
12 papers · 1 filter
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
Thilo von Neumann, Christoph Boeddeker, Marc Delcroix +1
The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker spe…
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
Reinhold Haeb-Umbach, Tomohiro Nakatani, Marc Delcroix +2
Multi-channel acoustic signal processing is a well-established and powerful tool to exploit the spatial diversity between a target signal and non-target or noise sources for signal…
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
Peter Vieting, Simon Berger, Thilo von Neumann +3
Many real-life applications of automatic speech recognition (ASR) require processing of overlapped speech. A common method involves first separating the speech into overlap-free st…
Post-Processing Independent Evaluation of Sound Event Detection Systems
Janek Ebbers, Reinhold Haeb-Umbach, Romain Serizel
Due to the high variation in the application requirements of sound event detection (SED) systems, it is not sufficient to evaluate systems only in a single operating mode. Therefor…
A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures
Tobias Cord-Landwehr, Christoph Boeddeker, Cătălin Zorilă +2
We introduce a monaural neural speaker embeddings extractor that computes an embedding for each speaker present in a speech mixture. To allow for supervised training, a teacher-stu…
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
Christoph Boeddeker, Aswin Shanmugam Subramanian, Gordon Wichern +2
Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-spea…