4 papers
Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition
Jingjing Xu, Zijian Yang, Mohammad Zeineldeen +3
BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates ps…
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
Eugen Beck, Sarah Beranek, Uma Moothiringote +4
Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into short segments, consist of rea…
Dynamic Acoustic Model Architecture Optimization in Training for ASR
Jingjing Xu, Zijian Yang, Albert Zeyer +3
Architecture design is inherently complex. Existing approaches rely on either handcrafted rules, which demand extensive empirical expertise, or automated methods like neural archit…
Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression
Jingjing Xu, Eugen Beck, Zijian Yang +1
ASR systems are deployed across diverse environments, each with specific hardware constraints. We use supernet training to jointly train multiple encoders of varying sizes, enablin…