3 papers
cs.LG2024
A Comprehensive Solution to Connect Speech Encoder and Large Language Model for ASR
Van Tung Pham, Yist Lin, Tao Han +4
Recent works have shown promising results in connecting speech encoders to large language models (LLMs) for speech recognition. However, several limitations persist, including limi…
cs.CL2023
Text-only Domain Adaptation using Unified Speech-Text Representation in Transducer
Lu Huang, Boyu Li, Jun Zhang +2
Domain adaptation using text-only corpus is challenging in end-to-end(E2E) speech recognition. Adaptation by synthesizing audio from text through TTS is resource-consuming. We pres…
cs.CL2023
CIF-PT: Bridging Speech and Text Representations for Spoken Language Understanding via Continuous Integrate-and-Fire Pre-Training
Linhao Dong, Zhecheng An, Peihao Wu +3
Speech or text representation generated by pre-trained models contains modal-specific information that could be combined for benefiting spoken language understanding (SLU) tasks. I…