1 paper
Mingwei Xu, Xuan Lin, Xinnan Guo +2
While linear attention reduces the quadratic complexity of standard Transformers to linear time, it often lags behind in expressivity due to the removal of softmax normalization. T…