29 citations · 53 across the 11 of their papers we have counts for
14 papers
Improved Speech Pre-Training with Supervision-Enhanced Acoustic Unit
Pengcheng Li, Genshun Wan, Fenglin Ding +4
Speech pre-training has shown great success in learning useful and general latent representations from large-scale unlabeled data. Based on a well-designed self-supervised learning…
Progressive Multi-Scale Self-Supervised Learning for Speech Recognition
Genshun Wan, Tan Liu, Hang Chen +3
Self-supervised learning (SSL) models have achieved considerable improvements in automatic speech recognition (ASR). In addition, ASR performance could be further improved if the m…
Improved Self-Supervised Multilingual Speech Representation Learning Combined with Auxiliary Language Information
Fenglin Ding, Genshun Wan, Pengcheng Li +2
Multilingual end-to-end models have shown great improvement over monolingual systems. With the development of pre-training methods on speech, self-supervised multilingual speech re…
Self-Supervised Audio-Visual Speech Representations Learning By Multimodal Self-Distillation
Jing-Xuan Zhang, Genshun Wan, Zhen-Hua Ling +3
In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module,…
Siamese Object Tracking for Vision-Based UAM Approaching with Pairwise Scale-Channel Attention
Guangze Zheng, Changhong Fu, Junjie Ye +3
Although the manipulating of the unmanned aerial manipulator (UAM) has been widely studied, vision-based UAM approaching, which is crucial to the subsequent manipulating, generally…
Adaptive Network Combination for Single-Image Reflection Removal: A Domain Generalization Perspective
Ming Liu, Jianan Pan, Zifei Yan +2
Recently, multiple synthetic and real-world datasets have been built to facilitate the training of deep single image reflection removal (SIRR) models. Meanwhile, diverse testing se…