4 citations · 4 across the 1 of their papers we have counts for
3 papers
cs.CV2023
Bootstrap Masked Visual Modeling via Hard Patches Mining
Haochen Wang, Junsong Fan, Yuxi Wang +4
Masked visual modeling has attracted much attention due to its promising potential in learning generalizable representations. Typical approaches urge models to predict specific con…
cs.CV2023
Semantic-Aware Autoregressive Image Modeling for Visual Representation Learning
Kaiyou Song, Shan Zhang, Tong Wang
The development of autoregressive modeling (AM) in computer vision lags behind natural language processing (NLP) in self-supervised pre-training. This is mainly caused by the chall…
cs.CV2023★ 4 cited
DropPos: Pre-Training Vision Transformers by Reconstructing Dropped Positions
Haochen Wang, Junsong Fan, Yuxi Wang +3
As it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enha…