activity
20172023
most citedTowards Realistic Visual Dubbing with Heterogeneous Sources

34 citations · 114 across the 28 of their papers we have counts for

collaborators
Showing eess.ASShow all

9 papers · 1 filter

eess.AS2022

Language Adaptive Cross-lingual Speech Representation Learning with Sparse Sharing Sub-networks

Yizhou Lu, Mingkun Huang, Xinghua Qu +2

Unsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across mult…

eess.AS20225 cited

Improving Non-native Word-level Pronunciation Scoring with Phone-level Mixup Data Augmentation and Multi-source Information

Kaiqi Fu, Shaojun Gao, Kai Wang +3

Deep learning-based pronunciation scoring models highly rely on the availability of the annotated non-native data, which is costly and has scalability issues. To deal with the data…

eess.AS20221 cited

S3T: Self-Supervised Pre-training with Swin Transformer for Music Classification

Hang Zhao, Chen Zhang, Belei Zhu +2

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive e…

eess.AS202111 cited

Cross-speaker Emotion Transfer Based on Speaker Condition Layer Normalization and Semi-Supervised Training in Text-To-Speech

Pengfei Wu, Junjie Pan, Chenchang Xu +4

In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due t…

eess.AS20213 cited

Improving Pseudo-label Training For End-to-end Speech Recognition Using Gradient Mask

Shaoshi Ling, Chen Shen, Meng Cai +1

In the recent trend of semi-supervised speech recognition, both self-supervised representation learning and pseudo-labeling have shown promising results. In this paper, we propose…

eess.AS2021

HMM-Free Encoder Pre-Training for Streaming RNN Transducer

Lu Huang, Jingyu Sun, Yufeng Tang +4

This work describes an encoder pre-training procedure using frame-wise label to improve the training of streaming recurrent neural network transducer (RNN-T) model. Streaming RNN-T…