1 paper · 1 filter
Yifan Zhu, Yekai Pan, Chen Ding
High-performance attention kernels are essential for Large Language Models. This paper presents analysis of CuTile-based Flash Attention memory behavior and a technique to improve…