2 papers
cs.LG2026
KVBuffer: IO-aware Serving for Linear Attention
Longwei Zou, Lin Zhong
Linear attention has recently gained significant attention for long-context inference due to its constant decoding cost with respect to context length. However, existing serving sy…
cs.CL2025
InstCache: A Predictive Cache for LLM Serving
Longwei Zou, Yan Liu, Jiamu Kang +3
The revolutionary capabilities of Large Language Models (LLMs) are attracting rapidly growing popularity and leading to soaring user requests to inference serving systems. Caching…