3 papers
eess.AS2024
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
Zhihong Lei, Xingyu Na, Mingbin Xu +5
Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a…
cs.LG2024
Focused Discriminative Training For Streaming CTC-Trained Automatic Speech Recognition Models
Adnan Haider, Xingyu Na, Erik McDermott +3
This paper introduces a novel training framework called Focused Discriminative Training (FDT) to further improve streaming word-piece end-to-end (E2E) automatic speech recognition…
eess.AS2024
Enhancing CTC-based speech recognition with diverse modeling units
Shiyi Han, Zhihong Lei, Mingbin Xu +2
In recent years, the evolution of end-to-end (E2E) automatic speech recognition (ASR) models has been remarkable, largely due to advances in deep learning architectures like transf…