18 citations · 79 across the 52 of their papers we have counts for
5 papers · 2 filters
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition
Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter
Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been sho…
Investigating the Effect of Language Models in Sequence Discriminative Training for Neural Transducers
Zijian Yang, Wei Zhou, Ralf Schlüter +1
In this work, we investigate the effect of language models (LMs) with different context lengths and label units (phoneme vs. word) used in sequence discriminative training for phon…
Mixture Encoder for Joint Speech Separation and Recognition
Simon Berger, Peter Vieting, Christoph Boeddeker +2
Multi-speaker automatic speech recognition (ASR) is crucial for many real-world applications, but it requires dedicated modeling techniques. Existing approaches can be divided into…
RASR2: The RWTH ASR Toolkit for Generic Sequence-to-sequence Speech Recognition
Wei Zhou, Eugen Beck, Simon Berger +2
Modern public ASR tools usually provide rich support for training various sequence-to-sequence (S2S) models, but rather simple support for decoding open-vocabulary scenarios only.…
Analyzing And Improving Neural Speaker Embeddings for ASR
Christoph Lüscher, Jingjing Xu, Mohammad Zeineldeen +2
Neural speaker embeddings encode the speaker's speech characteristics through a DNN model and are prevalent for speaker verification tasks. However, few studies have investigated t…