Showing 2026Show all
2 papers · 1 filter
cs.NI2026
Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty
Jiaming Cheng, Duong The Do, Duong Tung Nguyen
KV cache memory is a primary bottleneck in modern LLM serving systems deployed on GPU clusters. A fundamental challenge is that the KV cache must be reserved upon request arrival,…
cs.LG2026
Scalable Joint Resource Allocation for SLO-Constrained LLM Inference in Heterogeneous GPU Clouds
Jiaming Cheng, Duong Tung Nguyen
Serving large language model (LLM) inference in cloud environments requires jointly optimizing model selection, GPU provisioning, parallelism configuration, and workload routing un…