1 paper · 1 filter
Mutian He, Philip N. Garner
Linear-attention models that compress the entire input sequence into a fixed-size recurrent state offer an efficient alternative to Transformers, but their finite memory induces fo…