1 paper
James O'Neill, Robert Clancy, Mariia Matskevichus +1
The key-value (KV) cache is a primary memory bottleneck in Transformers. We propose Low-Rank Key-Value (LRKV) attention, which reduces KV cache memory by exploiting redundancy acro…