activity
20152022
most citedLarge Graph Analysis in the GMine System

29 citations · 53 across the 11 of their papers we have counts for

collaborators

14 papers

eess.AS20221 cited

Improved Speech Pre-Training with Supervision-Enhanced Acoustic Unit

Pengcheng Li, Genshun Wan, Fenglin Ding +4

Speech pre-training has shown great success in learning useful and general latent representations from large-scale unlabeled data. Based on a well-designed self-supervised learning…

eess.AS2022

Progressive Multi-Scale Self-Supervised Learning for Speech Recognition

Genshun Wan, Tan Liu, Hang Chen +3

Self-supervised learning (SSL) models have achieved considerable improvements in automatic speech recognition (ASR). In addition, ASR performance could be further improved if the m…

eess.AS2022

Improved Self-Supervised Multilingual Speech Representation Learning Combined with Auxiliary Language Information

Fenglin Ding, Genshun Wan, Pengcheng Li +2

Multilingual end-to-end models have shown great improvement over monolingual systems. With the development of pre-training methods on speech, self-supervised multilingual speech re…

eess.AS2022

Self-Supervised Audio-Visual Speech Representations Learning By Multimodal Self-Distillation

Jing-Xuan Zhang, Genshun Wan, Zhen-Hua Ling +3

In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module,…

cs.CV2022

Siamese Object Tracking for Vision-Based UAM Approaching with Pairwise Scale-Channel Attention

Guangze Zheng, Changhong Fu, Junjie Ye +3

Although the manipulating of the unmanned aerial manipulator (UAM) has been widely studied, vision-based UAM approaching, which is crucial to the subsequent manipulating, generally…

cs.CV20222 cited

Adaptive Network Combination for Single-Image Reflection Removal: A Domain Generalization Perspective

Ming Liu, Jianan Pan, Zifei Yan +2

Recently, multiple synthetic and real-world datasets have been built to facilitate the training of deep single image reflection removal (SIRR) models. Meanwhile, diverse testing se…