1 paper
Jiaxuan Chen, Jianshu She, Ye Yuan +5
LLM serving systems are provisioned for peak load to meet strict latency targets, leaving substantial GPU compute idle whenever traffic falls below peak. We present DeltaServe, a h…