4 papers
Service-Induced Congestion in Memory-Constrained LLM Serving
Ruicheng Ao, Jing Dong, Gan Luo +1
In large language model (LLM) serving, each request accumulates persistent graphics processing unit (GPU) memory during service as its key-value cache grows with every generated to…
Data-Driven Stochastic Modeling Using Autoregressive Sequence Models: Translating Event Tables to Queueing Dynamics
Daksh Mittal, Shunri Zheng, Jing Dong +1
While queueing network models are powerful tools for analyzing service systems, they traditionally require substantial human effort and domain expertise to construct. To make this…
QGym: Scalable Simulation and Benchmarking of Queuing Network Controllers
Haozhe Chen, Ang Li, Ethan Che +3
Queuing network control determines the allocation of scarce resources to manage congestion, a fundamental problem in manufacturing, communications, and healthcare. Compared to stan…
Stochastic Gradient Descent with Adaptive Data
Ethan Che, Jing Dong, Xin T. Tong
Stochastic gradient descent (SGD) is a powerful optimization technique that is particularly useful in online learning scenarios. Its convergence analysis is relatively well underst…