1 paper
Cong Liu, Xiaofang Li, Simon X. Yang
Vision Transformers (ViTs) commonly rely on injected positional mechanisms to address self-attention's permutation invariance. Motivated by the spatial regularities of natural imag…