1 paper
Minchul Kang, Changyong Shin, Jinwoo Jeong +6
Long-context LLM serving requires offloading KV caches to host-memory and SSDs, but existing mechanisms are not designed for such long contexts. We observe significant inefficienci…