4 papers
Neural domain alignment for spoken language recognition based on optimal transport
Xugang Lu, Peng Shen, Yu Tsao +1
Domain shift poses a significant challenge in cross-domain spoken language recognition (SLR) by reducing its effectiveness. Unsupervised domain adaptation (UDA) algorithms have bee…
Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-based ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for…
Cross-modal Alignment with Optimal Transport for CTC-based ASR
Xugang Lu, Peng Shen, Yu Tsao +1
Temporal connectionist temporal classification (CTC)-based automatic speech recognition (ASR) is one of the most successful end to end (E2E) ASR frameworks. However, due to the tok…
Pronunciation-aware unique character encoding for RNN Transducer-based Mandarin speech recognition
Peng Shen, Xugang Lu, Hisashi Kawai
For Mandarin end-to-end (E2E) automatic speech recognition (ASR) tasks, compared to character-based modeling units, pronunciation-based modeling units could improve the sharing of…