collaborators

5 papers

cs.SD2026

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

Xugang Lu, Peng Shen, Yu Tsao +1

Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary…

cs.SD2026

Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

Hao Shi, Yuan Gao, Xugang Lu +1

Large Language Models (LLMs) are effective decoders for Serialized Output Training (SOT) in two-talker automatic speech recognition (ASR), but their performance degrades substantia…

cs.CL2026

New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASR

Xugang Lu, Peng Shen, Hisashi Kawai

Aligning acoustic and linguistic representations is a central challenge to bridge the pre-trained models in knowledge transfer for automatic speech recognition (ASR). This alignmen…

eess.AS2025

Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR

Xugang Lu, Peng Shen, Yu Tsao +1

Transferring linguistic knowledge from a pretrained language model (PLM) to acoustic feature learning has proven effective in enhancing end-to-end automatic speech recognition (E2E…

cs.CL2025

Retrieval-Augmented Speech Recognition Approach for Domain Challenges

Peng Shen, Xugang Lu, Hisashi Kawai

Speech recognition systems often face challenges due to domain mismatch, particularly in real-world applications where domain-specific data is unavailable because of data accessibi…