1 paper · 1 filter
Jiaming Cheng, Duong Tung Nguyen
This paper investigates the optimal allocation of large language model (LLM) inference workloads across heterogeneous edge data centers over time. Each data center features on-site…