1 paper
Qinwen Wang, Jieping Luo, Aoxiang Qin +5
Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet their fixed-capacity recurrent states li…