13 citations · 13 across the 6 of their papers we have counts for
1 paper · 1 filter
Minchul Kang, Changyong Shin, Jinwoo Jeong +6
Long-context LLM serving requires offloading KV caches to host-memory and SSDs, but existing mechanisms are not designed for such long contexts. We observe significant inefficienci…