Showing cs.NIShow all
3 papers · 1 filter
cs.NI2026
Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty
Jiaming Cheng, Duong The Do, Duong Tung Nguyen
KV cache memory is a primary bottleneck in modern LLM serving systems deployed on GPU clusters. A fundamental challenge is that the KV cache must be reserved upon request arrival,…
cs.NI2026
Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference
Jiaming Cheng, Duong Tung Nguyen
This paper investigates the optimal allocation of large language model (LLM) inference workloads across heterogeneous edge data centers over time. Each data center features on-site…
cs.NI2025
Delay-Aware Robust Edge Network Hardening Under Decision-Dependent Uncertainty
Jiaming Cheng, Duong Thuy Anh Nguyen, Ni Trieu +1
Edge computing promises to offer low-latency and ubiquitous computation to numerous devices at the network edge. For delay-sensitive applications, link delays can have a direct imp…