2 citations · 4 across the 10 of their papers we have counts for
4 papers · 1 filter
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
Xugang Lu, Peng Shen, Yu Tsao +1
Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary…
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Transferring linguistic knowledge from a pretrained language model (PLM) to an acoustic model has been shown to greatly improve the performance of automatic speech recognition (ASR…
Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition
Peng Shen, Xugang Lu, Hisashi Kawai
Multi-talker overlapped speech recognition remains a significant challenge, requiring not only speech recognition but also speaker diarization tasks to be addressed. In this paper,…
Cross-scale Attention Model for Acoustic Event Classification
Xugang Lu, Peng Shen, Sheng Li +2
A major advantage of a deep convolutional neural network (CNN) is that the focused receptive field size is increased by stacking multiple convolutional layers. Accordingly, the mod…