1 paper
Mutian He, Philip N. Garner
Linear-attention models that compress the entire input sequence into a fixed-size recurrent state offer an efficient alternative to Transformers, but their finite memory induces fo…