collaborators

5 papers

eess.AS2025

Segmental Attention Decoding With Long Form Acoustic Encodings

Pawel Swietojanski, Xinwei Li, Mingbin Xu +3

We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to en…

cs.CL2025

Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition

Takaaki Hori, Martin Kocour, Adnan Haider +2

This paper presents an efficient decoding approach for end-to-end automatic speech recognition (E2E-ASR) with large language models (LLMs). Although shallow fusion is the most comm…

eess.AS2024

Optimizing Contextual Speech Recognition Using Vector Quantization for Efficient Retrieval

Nikolaos Flemotomos, Roger Hsiao, Pawel Swietojanski +3

Neural contextual biasing allows speech recognition models to leverage contextually relevant information, leading to improved transcription accuracy. However, the biasing mechanism…

eess.AS2024

Optimizing Byte-level Representation for End-to-end ASR

Roger Hsiao, Liuhui Deng, Erik McDermott +2

We propose a novel approach to optimizing a byte-level representation for end-to-end automatic speech recognition (ASR). Byte-level representation is often used by large scale mult…

cs.LG2024

Focused Discriminative Training For Streaming CTC-Trained Automatic Speech Recognition Models

Adnan Haider, Xingyu Na, Erik McDermott +3

This paper introduces a novel training framework called Focused Discriminative Training (FDT) to further improve streaming word-piece end-to-end (E2E) automatic speech recognition…