1 paper
Congpei Qiu, Zhaoyu Hu, Wei Ke +3
Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Pri…