26 citations · 60 across the 8 of their papers we have counts for
13 papers
When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
Guangting Wang, Yucheng Zhao, Chuanxin Tang +2
Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. Howe…
Self-Supervised Visual Representations Learning by Contrastive Mask Prediction
Yucheng Zhao, Guangting Wang, Chong Luo +2
Advanced self-supervised visual representation learning methods rely on the instance discrimination (ID) pretext task. We point out that the ID task has an implicit semantic consis…
Unsupervised Visual Representation Learning by Tracking Patches in Video
Guangting Wang, Yizhou Zhou, Chong Luo +3
Inspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to…
General-Purpose Speech Representation Learning through a Self-Supervised Multi-Granularity Framework
Yucheng Zhao, Dacheng Yin, Chong Luo +4
This paper presents a self-supervised learning framework, named MGF, for general-purpose speech representation learning. In the design of MGF, speech hierarchy is taken into consid…
VAE^2: Preventing Posterior Collapse of Variational Video Predictions in the Wild
Yizhou Zhou, Chong Luo, Xiaoyan Sun +2
Predicting future frames of video sequences is challenging due to the complex and stochastic nature of the problem. Video prediction methods based on variational auto-encoders (VAE…
Online Speaker Diarization with Relation Network
Xiang Li, Yucheng Zhao, Chong Luo +1
In this paper, we propose an online speaker diarization system based on Relation Network, named RenoSD. Unlike conventional diariztion systems which consist of several independentl…