34 citations · 114 across the 28 of their papers we have counts for
9 papers · 1 filter
Language Adaptive Cross-lingual Speech Representation Learning with Sparse Sharing Sub-networks
Yizhou Lu, Mingkun Huang, Xinghua Qu +2
Unsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across mult…
Improving Non-native Word-level Pronunciation Scoring with Phone-level Mixup Data Augmentation and Multi-source Information
Kaiqi Fu, Shaojun Gao, Kai Wang +3
Deep learning-based pronunciation scoring models highly rely on the availability of the annotated non-native data, which is costly and has scalability issues. To deal with the data…
S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification
Hang Zhao, Chen Zhang, Belei Zhu +2
In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive e…
Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech
Pengfei Wu, Junjie Pan, Chenchang Xu +4
In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due t…
Improving Pseudo-label Training For End-to-end Speech Recognition Using Gradient Mask
Shaoshi Ling, Chen Shen, Meng Cai +1
In the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose…
HMM-Free Encoder Pre-Training for Streaming RNN Transducer
Lu Huang, Jingyu Sun, Yufeng Tang +4
This work describes an encoder pre-training procedure using frame-wise label to improve the training of streaming recurrent neural network transducer (RNN-T) model. Streaming RNN-T…