1 paper
Yichun Xu, Navjot K. Khaira, Tejinder Singh
The key-value (KV) cache is a foundational optimization in Transformer-based large language models (LLMs), eliminating redundant recomputation of past token representations during…