Showing eess.ASShow all
3 papers · 1 filter
eess.AS2025
Segmental Attention Decoding With Long Form Acoustic Encodings
Pawel Swietojanski, Xinwei Li, Mingbin Xu +3
We address the fundamental incompatibility of attention-based encoder-decoder (AED) models with long-form acoustic encodings. AED models trained on segmented utterances learn to en…
eess.AS2024
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
Zhihong Lei, Xingyu Na, Mingbin Xu +5
Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a…
eess.AS2024
Enhancing CTC-based speech recognition with diverse modeling units
Shiyi Han, Zhihong Lei, Mingbin Xu +2
In recent years, the evolution of end-to-end (E2E) automatic speech recognition (ASR) models has been remarkable, largely due to advances in deep learning architectures like transf…