5 papers · 1 filter
Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition
Jingjing Xu, Zijian Yang, Mohammad Zeineldeen +3
BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates ps…
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
Zijian Yang, Jörg Barkoczi, Ralf Schlüter +1
Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how…
A Comparative Analysis on ASR System Combination for Attention, CTC, Factored Hybrid, and Transducer Models
Noureldin Bayoumi, Robin Schmitt, Tina Raissi +3
Combination approaches for speech recognition (ASR) systems cover structured sentence-level or word-based merging techniques as well as combination of model scores during beam sear…
Label-Context-Dependent Internal Language Model Estimation for CTC
Zijian Yang, Minh-Nghia Phan, Ralf Schlüter +1
Although connectionist temporal classification (CTC) has the label context independence assumption, it can still implicitly learn a context-dependent internal language model (ILM)…
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
Tina Raissi, Ralf Schlüter, Hermann Ney
Current time-synchronous sequence-to-sequence automatic speech recognition (ASR) models are trained by using sequence level cross-entropy that sums over all alignments. Due to the…