activity
20242026
collaborators

5 papers

cs.NI2026

Robust KV Cache Management for LLM Serving under Output Token Length Uncertainty

Jiaming Cheng, Duong The Do, Duong Tung Nguyen

KV cache memory is a primary bottleneck in modern LLM serving systems deployed on GPU clusters. A fundamental challenge is that the KV cache must be reserved upon request arrival,…

cs.LG2026

Scalable Joint Resource Allocation for SLO-Constrained LLM Inference in Heterogeneous GPU Clouds

Jiaming Cheng, Duong Tung Nguyen

Serving large language model (LLM) inference in cloud environments requires jointly optimizing model selection, GPU provisioning, parallelism configuration, and workload routing un…

cs.NI2026

Green-LLM: Optimal Workload Allocation for Environmentally-Aware Distributed Inference

Jiaming Cheng, Duong Tung Nguyen

This paper investigates the optimal allocation of large language model (LLM) inference workloads across heterogeneous edge data centers over time. Each data center features on-site…

cs.NI2025

Delay-Aware Robust Edge Network Hardening Under Decision-Dependent Uncertainty

Jiaming Cheng, Duong Thuy Anh Nguyen, Ni Trieu +1

Edge computing promises to offer low-latency and ubiquitous computation to numerous devices at the network edge. For delay-sensitive applications, link delays can have a direct imp…

math.OC2024

Robust Dynamic Edge Service Placement Under Spatio-Temporal Correlated Demand Uncertainty

Jiaming Cheng, Duong Thuy Anh Nguyen, Duong Tung Nguyen

Edge computing allows Service Providers (SPs) to enhance user experience by placing their services closer to the network edge. Determining the optimal provisioning of edge resource…