5 papers
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
Xugang Lu, Peng Shen, Yu Tsao +1
Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary…
New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR
Xugang Lu, Peng Shen, Hisashi Kawai
Aligning acoustic and linguistic representations is a central challenge to bridge the pre-trained models in knowledge transfer for automatic speech recognition (ASR). This alignmen…
Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Transferring linguistic knowledge from a pretrained language model (PLM) to acoustic feature learning has proven effective in enhancing end-to-end automatic speech recognition (E2E…
Linguistic Knowledge Transfer Learning for Speech Enhancement
Kuo-Hsuan Hung, Xugang Lu, Szu-Wei Fu +4
Linguistic knowledge plays a crucial role in spoken language comprehension. It provides essential semantic and syntactic context for speech perception in noisy environments. Howeve…
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR…