2 papers
cs.AI2026
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
William Nixon, Jon Durbin, Florian Standhartinger +2
Large Language Model (LLM) serving has become a critical cloud workload, and realistic traces are essential for motivating and benchmarking serving systems. However, existing LLM s…
cs.DC2026
Learning-Augmented Heuristics: Simple, yet Smart, Robust and Interpretable Cache Eviction
Haocheng Xia, William Nixon, Bintang Dwi Marthen +2
Caching is widely used across the system stack to improve performance and efficiency, with eviction algorithms at its core. Existing cache eviction policies fall into two broad cat…