activity
20202023
most citedVideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

441 citations · 875 across the 13 of their papers we have counts for

collaborators
Showing 2023Show all

9 papers · 1 filter

cs.CV2023

Bootstrapping SparseFormers from Vision Foundation Models

Ziteng Gao, Zhan Tong, Kevin Qinghong Lin +2

The recently proposed SparseFormer architecture provides an alternative approach to visual understanding by utilizing a significantly lower number of visual tokens via adjusting Ro…

cs.CV2023★ 12 cited

Advancing Vision Transformers with Group-Mix Attention

Chongjian Ge, Xiaohan Ding, Zhan Tong +4

Vision Transformers (ViTs) have been shown to enhance visual recognition through modeling long-range dependencies with multi-head self-attention (MHSA), which is typically formulat…

cs.CV2023

Speed Co-Augmentation for Unsupervised Audio-Visual Pre-training

Jiangliu Wang, Jianbo Jiao, Yibing Song +5

This work aims to improve unsupervised audio-visual pre-training. Inspired by the efficacy of data augmentation in visual contrastive learning, we propose a novel speed co-augmenta…

cs.CV2023★ 2 cited

TVTSv2: Learning Out-of-the-box Spatiotemporal Visual Representations at Scale

Ziyun Zeng, Yixiao Ge, Zhan Tong +3

The ultimate goal for foundation models is realizing task-agnostic, i.e., supporting out-of-the-box usage without task-specific fine-tuning. Although breakthroughs have been made i…

cs.CV2023★ 15 cited

VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Limin Wang, Bingkun Huang, Zhiyu Zhao +5

Scale is the primary factor for building a powerful foundation model that could well generalize to a variety of downstream tasks. However, it is still challenging to train video fo…

cs.CV2023★ 6 cited

SparseFormer: Sparse Visual Recognition via Limited Latent Tokens

Ziteng Gao, Zhan Tong, Limin Wang +1

Human visual recognition is a sparse process, where only a few salient visual cues are attended to rather than traversing every detail uniformly. However, most current vision netwo…