7 citations · 56 across the 64 of their papers we have counts for
24 papers
Speech2Video: Cross-Modal Distillation for Speech to Video Generation
Shijing Si, Jianzong Wang, Xiaoyang Qu +4
This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertain…
Variational Information Bottleneck for Effective Low-resource Audio Classification
Shijing Si, Jianzong Wang, Huiming Sun +6
Large-scale deep neural networks (DNNs) such as convolutional neural networks (CNNs) have achieved impressive performance in audio classification for their powerful capacity and st…
Federated Learning with Dynamic Transformer for Text to Speech
Zhenhou Hong, Jianzong Wang, Xiaoyang Qu +3
Text to speech (TTS) is a crucial task for user interaction, but TTS model training relies on a sizable set of high-quality original datasets. Due to privacy and security issues, t…
Loss Prediction: End-to-End Active Learning Approach For Speech Recognition
Jian Luo, Jianzong Wang, Ning Cheng +1
End-to-end speech recognition systems usually require huge amounts of labeling resource, while annotating the speech data is complicated and expensive. Active learning is the solut…
Dropout Regularization for Self-Supervised Learning of Transformer Encoder Speech Representation
Jian Luo, Jianzong Wang, Ning Cheng +1
Predicting the altered acoustic frames is an effective way of self-supervised learning for speech representation. However, it is challenging to prevent the pretrained model from ov…
Efficient Client Contribution Evaluation for Horizontal Federated Learning
Jie Zhao, Xinghua Zhu, Jianzong Wang +1
In federated learning (FL), fair and accurate measurement of the contribution of each federated participant is of great significance. The level of contribution not only provides a…