4 citations · 4 across the 1 of their papers we have counts for
1 paper
Guoxi Huang, Hongtao Fu, Adrian G. Bors
Deeper Vision Transformers (ViTs) are more challenging to train. We expose a degradation problem in deeper layers of ViT when using masked image modeling (MIM) for pre-training. To…