58 citations · 97 across the 38 of their papers we have counts for
14 papers · 1 filter
Boosting Star-GANs for Voice Conversion with Contrastive Discriminator
Shijing Si, Jianzong Wang, Xulong Zhang +3
Nonparallel multi-domain voice conversion methods such as the StarGAN-VCs have been widely applied in many scenarios. However, the training of these models usually poses a challeng…
Loss Prediction: End-to-End Active Learning Approach For Speech Recognition
Jian Luo, Jianzong Wang, Ning Cheng +1
End-to-end speech recognition systems usually require huge amounts of labeling resource, while annotating the speech data is complicated and expensive. Active learning is the solut…
Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech Representation
Jian Luo, Jianzong Wang, Ning Cheng +1
Predicting the altered acoustic frames is an effective way of self-supervised learning for speech representation. However, it is challenging to prevent the pretrained model from ov…
Unidirectional Memory-Self-Attention Transducer for Online Speech Recognition
Jian Luo, Jianzong Wang, Ning Cheng +1
Self-attention models have been successfully applied in end-to-end speech recognition systems, which greatly improve the performance of recognition accuracy. However, such attentio…
LVCNet: Efficient Condition-Dependent Modeling Network for Waveform Generation
Zhen Zeng, Jianzong Wang, Ning Cheng +1
In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use o…
GraphPB: Graphical Representations of Prosody Boundary in Speech Synthesis
Aolan Sun, Jianzong Wang, Ning Cheng +4
This paper introduces a graphical representation approach of prosody boundary (GraphPB) in the task of Chinese speech synthesis, intending to parse the semantic and syntactic relat…