1 paper · 1 filter
Zhongchun Zhou, Chengtao Lai, Songtao Mao
In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V…