1 paper · 1 filter
Runchao Li, Yao Fu, Mu Sheng +3
The efficacy of Large Language Models (LLMs) in long-context tasks is often hampered by the substantial memory footprint and computational demands of the Key-Value (KV) cache. Curr…