activity
20242026
collaborators

5 papers

cs.SD2026

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

Xugang Lu, Peng Shen, Yu Tsao +1

Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary…

cs.CL2026

New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR

Xugang Lu, Peng Shen, Hisashi Kawai

Aligning acoustic and linguistic representations is a central challenge to bridge the pre-trained models in knowledge transfer for automatic speech recognition (ASR). This alignmen…

eess.AS2025

Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR

Xugang Lu, Peng Shen, Yu Tsao +1

Transferring linguistic knowledge from a pretrained language model (PLM) to acoustic feature learning has proven effective in enhancing end-to-end automatic speech recognition (E2E…

cs.CL2025

Linguistic Knowledge Transfer Learning for Speech Enhancement

Kuo-Hsuan Hung, Xugang Lu, Szu-Wei Fu +4

Linguistic knowledge plays a crucial role in spoken language comprehension. It provides essential semantic and syntactic context for speech perception in noisy environments. Howeve…

cs.SD2024

Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR

Xugang Lu, Peng Shen, Yu Tsao +1

Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR…