1 paper · 1 filter
Wenshuai Yao, Wenyong Zhou, Hanyong Shao +5
Top-K sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identifying this subset still re…