1 paper
Dong Le, Thong Nguyen, Cong-Duy Nguyen +1
Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tas…