activity
20202024
most citedApplying Wav2vec2.0 to Speech Recognition in Various Low-resource Languages

58 citations · 97 across the 38 of their papers we have counts for

collaborators
Showing eess.ASShow all

14 papers · 1 filter

eess.AS2022

Boosting Star-GANs for Voice Conversion with Contrastive Discriminator

Shijing Si, Jianzong Wang, Xulong Zhang +3

Nonparallel multi-domain voice conversion methods such as the StarGAN-VCs have been widely applied in many scenarios. However, the training of these models usually poses a challeng…

eess.AS2021

Loss Prediction: End-to-End Active Learning Approach For Speech Recognition

Jian Luo, Jianzong Wang, Ning Cheng +1

End-to-end speech recognition systems usually require huge amounts of labeling resource, while annotating the speech data is complicated and expensive. Active learning is the solut…

eess.AS2021

Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech Representation

Jian Luo, Jianzong Wang, Ning Cheng +1

Predicting the altered acoustic frames is an effective way of self-supervised learning for speech representation. However, it is challenging to prevent the pretrained model from ov…

eess.AS20212 cited

Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition

Jian Luo, Jianzong Wang, Ning Cheng +1

Self-attention models have been successfully applied in end-to-end speech recognition systems, which greatly improve the performance of recognition accuracy. However, such attentio…

eess.AS20214 cited

LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation

Zhen Zeng, Jianzong Wang, Ning Cheng +1

In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use o…

eess.AS20202 cited

GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis

Aolan Sun, Jianzong Wang, Ning Cheng +4

This paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relat…