7 papers
Refined Statistical Bounds for Classification Error Mismatches with Constrained Bayes Error
Zijian Yang, Vahe Eminyan, Ralf Schlüter +1
In statistical classification/multiple hypothesis testing and machine learning, a model distribution estimated from the training data is usually applied to replace the unknown true…
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition
Nick Rossenbach, Benedikt Hilmes, Ralf Schlüter
Synthetic data generated by text-to-speech (TTS) systems can be used to improve automatic speech recognition (ASR) systems in low-resource or domain mismatch tasks. It has been sho…
Investigating the Effect of Language Models in Sequence Discriminative Training for Neural Transducers
Zijian Yang, Wei Zhou, Ralf Schlüter +1
In this work, we investigate the effect of language models (LMs) with different context lengths and label units (phoneme vs. word) used in sequence discriminative training for phon…
End-to-End Training of a Neural HMM with Label and Transition Probabilities
Daniel Mann, Tina Raissi, Wilfried Michel +2
We investigate a novel modeling approach for end-to-end neural network training using hidden Markov models (HMM) where the transition probabilities between hidden states are modele…
Comparative Analysis of the wav2vec 2.0 Feature Extractor
Peter Vieting, Ralf Schlüter, Hermann Ney
Automatic speech recognition (ASR) systems typically use handcrafted feature extraction pipelines. To avoid their inherent information loss and to achieve more consistent modeling…
Improving the Training Recipe for a Robust Conformer-based Hybrid Model
Mohammad Zeineldeen, Jingjing Xu, Christoph Lüscher +2
Speaker adaptation is important to build robust automatic speech recognition (ASR) systems. In this work, we investigate various methods for speaker adaptive training (SAT) based o…