1 citations · 1 across the 3 of their papers we have counts for
3 papers
eess.AS2022★ 1 cited
Improved Speech Pre-Training with Supervision-Enhanced Acoustic Unit
Pengcheng Li, Genshun Wan, Fenglin Ding +4
Speech pre-training has shown great success in learning useful and general latent representations from large-scale unlabeled data. Based on a well-designed self-supervised learning…
eess.AS2022
Progressive Multi-Scale Self-Supervised Learning for Speech Recognition
Genshun Wan, Tan Liu, Hang Chen +3
Self-supervised learning (SSL) models have achieved considerable improvements in automatic speech recognition (ASR). In addition, ASR performance could be further improved if the m…
eess.AS2022
Self-Supervised Audio-Visual Speech Representations Learning By Multimodal Self-Distillation
Jing-Xuan Zhang, Genshun Wan, Zhen-Hua Ling +3
In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module,…