1 paper · 1 filter
Vaisakh Shaj, Cameron Barker, Aidan Scannell +3
State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but their linear state updates limi…