1 paper
Rithin Nagaraj, Rupa Laalasa Oruganti, Prerna Subhashchandra Kunder +1
The quadratic scaling of Transformer self-attention has driven the adoption of sub-quadratic Selective State Space Models (SSMs) like Mamba, which compress past context into a fixe…