2 papers
cs.DC2026
Coda: Exploiting Admission Flexibility for Coding-Agent Serving
Youhe Jiang, Fangcheng Fu, Binhang Yuan +5
Coding agents powered by large language models (LLMs) repeatedly alternate between model inference and tool calls, creating long-lived sessions with reusable key-value (KV) states…
cs.AI2026
ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
Zheyu Shen, Guanhua Wang, Dezhan Tu +5
Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefix…