1 paper · 1 filter
Zhirui Chen, Peiyang Liu, Ling Shao
As Large Language Models (LLMs) scale to support context windows exceeding one million tokens, the linear growth of Key-Value (KV) cache imposes severe memory capacity and bandwidt…