1 paper
Khalid Shaikh, Asmit Kumar Singh, Rebecca Christopher Dsouza +1
Large language model (LLM) inference is increasingly bottlenecked by the Key-Value (KV) cache, yet the fine-grained structure of attention-head activations remains poorly understoo…