2 papers
eess.AS2025
Segmental Attention Decoding With Long Form Acoustic Encodings
Pawel Swietojanski, Xinwei Li, Mingbin Xu +3
We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to en…
eess.AS2024
Optimizing Contextual Speech Recognition Using Vector Quantization for Efficient Retrieval
Nikolaos Flemotomos, Roger Hsiao, Pawel Swietojanski +3
Neural contextual biasing allows speech recognition models to leverage contextually relevant information, leading to improved transcription accuracy. However, the biasing mechanism…