4 citations · 5 across the 2 of their papers we have counts for
1 paper · 1 filter
Haochen Wang, Junsong Fan, Yuxi Wang +3
As it is empirically observed that Vision Transformers (ViTs) are quite insensitive to the order of input tokens, the need for an appropriate self-supervised pretext task that enha…