5 papers
Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition
Jingjing Xu, Zijian Yang, Mohammad Zeineldeen +3
BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates ps…
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
Zijian Yang, Jörg Barkoczi, Ralf Schlüter +1
Unsupervised speech recognition is a task of training a speech recognition model with unpaired data. To determine when and how unsupervised speech recognition can succeed, and how…
Dynamic Acoustic Model Architecture Optimization in Training for ASR
Jingjing Xu, Zijian Yang, Albert Zeyer +3
Architecture design is inherently complex. Existing approaches rely on either handcrafted rules, which demand extensive empirical expertise, or automated methods like neural archit…
Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression
Jingjing Xu, Eugen Beck, Zijian Yang +1
ASR systems are deployed across diverse environments, each with specific hardware constraints. We use supernet training to jointly train multiple encoders of varying sizes, enablin…
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
Jingjing Xu, Wei Zhou, Zijian Yang +2
Varying-size models are often required to deploy ASR systems under different hardware and/or application constraints such as memory and latency. To avoid redundant training and opt…