4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 4 cited
DropPos: Pre-Training Vision Transformers by Reconstructing Dropped Positions
Haochen Wang, Junsong Fan, Yuxi Wang +3
As it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enha…
cs.CV2023
Hard Patches Mining for Masked Image Modeling
Haochen Wang, Kaiyou Song, Junsong Fan +3
Masked image modeling (MIM) has attracted much research attention due to its promising potential for learning scalable visual representations. In typical approaches, models usually…