1 paper
Liming Lu, Kaixi Qiu, Jiayu Zhou +6
Despite the remarkable progress of Large Language Models (LLMs), the escalating memory footprint of the Key-Value (KV) cache remains a critical bottleneck for efficient inference.…