2 citations · 6 across the 13 of their papers we have counts for
Showing cs.OSShow all
2 papers · 1 filter
cs.OS2026
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
Jing Zou, Shangyu Wu, Hancong Duan +2
Efficiently serving Large Language Models (LLMs) with persistent Prefix Key-Value (KV) Cache is critical for applications like conversational search and multi-turn dialogue. Servin…
cs.OS2025★ 1 cited
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
Hongchao Du, Shangyu Wu, Arina Kharlamova +2
Large Language Models (LLMs) face challenges for on-device inference due to high memory demands. Traditional methods to reduce memory usage often compromise performance and lack ad…