1 paper
William Meng, Benjamin Lee, Hong Wang
KV cache offloading enables long-context LLM inference by storing caches in CPU DRAM, but PCIe bandwidth limitations create severe bottlenecks. In this paper, we develops an analyt…