1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yongyu Mu, Yuzhang Wu, Yuchun Fan +9
To enhance the efficiency of the attention mechanism within large language models (LLMs), previous works primarily compress the KV cache or group attention heads, while largely ove…