Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
2Mamba2Furious: Linear in Complexity, Competitive in Accuracy
Gabriel Mongaras, Eric C. Larson
Linear attention transformers have become a strong alternative to softmax attention due to their efficiency. However, linear attention tends to be less expressive and results in re…
cs.LG2025
On the Expressiveness of Softmax Attention: A Recurrent Neural Network Perspective
Gabriel Mongaras, Eric C. Larson
Since its introduction, softmax attention has become the backbone of modern transformer architectures due to its expressiveness and scalability across a wide range of tasks. Howeve…
cs.LG2024
Cottention: Linear Transformers With Cosine Attention
Gabriel Mongaras, Trevor Dohm, Eric C. Larson
Attention mechanisms, particularly softmax attention, have been instrumental in the success of transformer-based models such as GPT. However, the quadratic memory complexity of sof…