1 paper
Wei-Yuan Su, Ruijie Zhang, Zheng Zhang
Vision Transformers (ViTs) achieve state-of-the-art performance but suffer from the O(N2) complexity of self-attention, making inference costly for high-resolution inputs. To ad…