2 citations · 2 across the 1 of their papers we have counts for
1 paper
Jiantong Jiang, Peiyu Yang, Rui Zhang +1
Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV) cache, which stores KV tensors during aut…