2 papers
eess.AS2025
Segmental Attention Decoding With Long Form Acoustic Encodings
Pawel Swietojanski, Xinwei Li, Mingbin Xu +3
We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to en…
cs.CL2025
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
Takaaki Hori, Martin Kocour, Adnan Haider +2
This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most comm…