2 papers
cs.CL2026
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
Gradwell Dzikanyanga, Weihao Yang, Hao Huang +4
Key-value (KV) caching is critical for efficient inference in large language models (LLMs), yet its memory footprint scales linearly with context length, resulting in a severe scal…
cs.OS2026
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
Shihao Wang, Jiahao Chen, Yanqi Pan +7
The prefill stage of long-context Retrieval-Augmented Generation (RAG) is severely bottlenecked by computational overhead. To mitigate this, recent methods assemble pre-calculated…