1 paper
Yingfa Chen, Yutong Wu, Chenyang Song +5
Grouped-Query Attention (GQA) is a widely adopted strategy for reducing the computational cost of attention layers in large language models (LLMs). However, current GQA configurati…