26 citations · 60 across the 8 of their papers we have counts for
9 papers · 1 filter
When Shift Operation Meets Vision Transformer: An Extremely Simple Alternative to Attention Mechanism
Guangting Wang, Yucheng Zhao, Chuanxin Tang +2
Attention mechanism has been widely believed as the key to success of vision transformers (ViTs), since it provides a flexible and powerful way to model spatial relationships. Howe…
Self-Supervised Visual Representations Learning by Contrastive Mask Prediction
Yucheng Zhao, Guangting Wang, Chong Luo +2
Advanced self-supervised visual representation learning methods rely on the instance discrimination (ID) pretext task. We point out that the ID task has an implicit semantic consis…
Unsupervised Visual Representation Learning by Tracking Patches in Video
Guangting Wang, Yizhou Zhou, Chong Luo +3
Inspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to…
VAE^2: Preventing Posterior Collapse of Variational Video Predictions in the Wild
Yizhou Zhou, Chong Luo, Xiaoyan Sun +2
Predicting future frames of video sequences is challenging due to the complex and stochastic nature of the problem. Video prediction methods based on variational auto-encoders (VAE…
Spatiotemporal Fusion in 3D CNNs: A Probabilistic View
Yizhou Zhou, Xiaoyan Sun, Chong Luo +2
Despite the success in still image recognition, deep neural networks for spatiotemporal signal tasks (such as human action recognition in videos) still suffers from low efficacy an…
Tracking by Instance Detection: A Meta-Learning Approach
Guangting Wang, Chong Luo, Xiaoyan Sun +2
We consider the tracking problem as a special type of object detection problem, which we call instance detection. With proper initialization, a detector can be quickly converted in…