2 papers
cs.CL2024
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
Yang Yuhang, Peng Yizhou, Eng Siong Chng +1
The integration of large language models (LLMs) with pre-trained speech models has opened up new avenues in automatic speech recognition (ASR). While LLMs excel in multimodal under…
eess.AS2022
Speech-text based multi-modal training with bidirectional attention for improved speech recognition
Yuhang Yang, Haihua Xu, Hao Huang +2
To let the state-of-the-art end-to-end ASR model enjoy data efficiency, as well as much more unpaired text data by multi-modal training, one needs to address two problems: 1) the s…