1 paper · 1 filter
Jian Lin, Jiazhi Mi, Zicong Hong +5
Supporting long-context LLMs is challenging due to the substantial memory demands of the key-value (KV) cache. Existing offloading systems store the full cache in host memory and s…