Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Subquadratic Algorithms and Hardness for Attention with Any Temperature
Shreya Gupta, Boyang Huang, Barna Saha +2
Despite the popularity of the Transformer architecture, the standard algorithm for computing Attention suffers from quadratic time complexity in context length . Alman and Song…
cs.LG2024
The I/O Complexity of Attention, or How Optimal is Flash Attention?
Barna Saha, Christopher Ye
Self-attention is at the heart of the popular Transformer architecture, yet suffers from quadratic time and memory complexity. The breakthrough FlashAttention algorithm revealed I/…