1 paper
Zuxiong Tan, Will Wei-Jen Wang, Wei Shao +3
Autoregressive large language model (LLM) decoding re-reads a growing key-value (KV) cache at every step, so long-context attention is bound by graphics processing unit (GPU) memor…