1 paper
Tuowei Wang, Liyun Chu, Ruwen Fan +1
The key-value (KV) cache has become the dominant contributor to memory consumption in large language model (LLM) inference. Although offloading KVCache from GPU high-bandwidth memo…