1 paper
Kaixuan He, Song Chen, Yi Kang
Vision Transformers (ViTs) incur significant computational overhead due to the quadratic complexity of self-attention relative to the token sequence length. While existing token re…