1 paper
Yiming Bian, Joshua M. Akey
The scalability of long-context large language models is fundamentally limited by the quadratic memory cost of exact self-attention, which often leads to out-of-memory (OOM) failur…