1 paper
Jiaxu Liu, Yuhe Bai, Xiangyu Yin +1
Modern autoregressive models rely on attention, yet the Softmax full attention in Transformers scales quadratically with sequence length. Sliding Window Attention (SWA) achieves li…