1 paper
Vishesh Tripathi, Abhay Kumar, Ramsha Khan
The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with sequence length. Grouped-query attention (GQA) reduces this cos…