4 papers
LLMs and Speech: Integration vs. Combination
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen +2
In this work, we study different approaches to utilize large language models (LLMs) for automatic speech recognition (ASR). Specifically, we compare the tight integration of an aco…
Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition
Mohammad Zeineldeen, Albert Zeyer, Haoran Zhang +3
Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately…
Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition
Jingjing Xu, Zijian Yang, Mohammad Zeineldeen +3
BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates ps…
The Conformer Encoder May Reverse the Time Dimension
Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen +2
We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, Further investigation shows that t…