activity
20162026
most citedWhy does CTC result in peaky behavior?

18 citations · 79 across the 52 of their papers we have counts for

collaborators
Showing 2023 · cs.CLShow all

5 papers · 2 filters

cs.CL2023

On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition

Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter

Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been sho…

cs.CL2023

Investigating the Effect of Language Models in Sequence Discriminative Training for Neural Transducers

Zijian Yang, Wei Zhou, Ralf Schlüter +1

In this work, we investigate the effect of language models (LMs) with different context lengths and label units (phoneme vs. word) used in sequence discriminative training for phon…

cs.CL2023

Mixture Encoder for Joint Speech Separation and Recognition

Simon Berger, Peter Vieting, Christoph Boeddeker +2

Multi-speaker automatic speech recognition (ASR) is crucial for many real-world applications, but it requires dedicated modeling techniques. Existing approaches can be divided into…

cs.CL2023★ 4 cited

RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech Recognition

Wei Zhou, Eugen Beck, Simon Berger +2

Modern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.…

cs.CL2023

Analyzing And Improving Neural Speaker Embeddings for ASR

Christoph Lüscher, Jingjing Xu, Mohammad Zeineldeen +2

Neural speaker embeddings encode the speaker's speech characteristics through a DNN model and are prevalent for speaker verification tasks. However, few studies have investigated t…