4 papers
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
Takaaki Hori, Martin Kocour, Adnan Haider +2
This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most comm…
Focused Discriminative Training For Streaming CTC-Trained Automatic Speech Recognition Models
Adnan Haider, Xingyu Na, Erik McDermott +3
This paper introduces a novel training framework called Focused Discriminative Training (FDT) to further improve streaming word-piece end-to-end (E2E) automatic speech recognition…
Optimizing Byte-level Representation for End-to-end ASR
Roger Hsiao, Liuhui Deng, Erik McDermott +2
We propose a novel approach to optimizing a byte-level representation for end-to-end automatic speech recognition (ASR). Byte-level representation is often used by large scale mult…
Revisiting ASR Error Correction with Specialized Models
Zijin Gu, Tatiana Likhomanenko, He Bai +3
Language models play a central role in automatic speech recognition (ASR), yet most methods rely on text-only models unaware of ASR error patterns. Recently, large language models…