1 paper
Alexander Tian, Aditya Ghai, Sanjit Neelam +2
Block sparse attention is a hardware friendly way to alleviate the key-value (KV) cache read bottleneck in large language models (LLMs). However, it is not prevalent among leading…