1 paper · 1 filter
Luning Wang, Shiyao Li, Xuefei Ning +4
Large Language Models (LLMs) have been widely adopted to process long-context tasks. However, the large memory overhead of the key-value (KV) cache poses significant challenges in…