2 papers
cs.LG2026
MosaicKV: Serving Long-Context LLM with Dynamic Two-D KV Cache Compression
Sheng Qiang, Ruiwei Chen, Yinpeng Wu +5
Long-context LLM services now sustain prompts with hundreds of thousands to millions of tokens, making the key-value (KV) cache a first-order serving cost. Because the cache grows…
cs.OS2025
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
Mingcong Han, Weihang Shen, Rong Chen +2
Modern autonomous applications are increasingly utilizing multiple heterogeneous processors (XPUs) to accelerate different stages of algorithm modules. However, existing runtime sy…