collaborators

5 papers

cs.SD2024

Robust Channel Learning for Large-Scale Radio Speaker Verification

Wenhao Yang, Jianguo Wei, Wenhuan Lu +2

Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifyi…

eess.AS2023

Neural domain alignment for spoken language recognition based on optimal transport

Xugang Lu, Peng Shen, Yu Tsao +1

Domain shift poses a significant challenge in cross-domain spoken language recognition (SLR) by reducing its effectiveness. Unsupervised domain adaptation (UDA) algorithms have bee…

eess.AS2023

Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR

Xugang Lu, Peng Shen, Yu Tsao +1

Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for…

eess.AS2023

Cross-modal Alignment with Optimal Transport for CTC-based ASR

Xugang Lu, Peng Shen, Yu Tsao +1

Temporal connectionist temporal classification (CTC)-based automatic speech recognition (ASR) is one of the most successful end to end (E2E) ASR frameworks. However, due to the tok…

cs.CL2022

Pronunciation-aware unique character encoding for RNN Transducer-based Mandarin speech recognition

Peng Shen, Xugang Lu, Hisashi Kawai

For Mandarin end-to-end (E2E) automatic speech recognition (ASR) tasks, compared to character-based modeling units, pronunciation-based modeling units could improve the sharing of…