2 papers
cs.DC2026
PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response
Yueying Li, Jiayang Chen, Yuanfan Chen +5
AI inference clusters are increasingly constrained by instantaneous power, not just energy: grid operators condition new capacity on demand response, imposing time-varying power ca…
cs.LG2026
Beyond Prediction: Tail-Aware Scheduling for LLM Inference
Yueying Li, Yuanfan Chen, Jiayang Chen +6
LLM serving exhibits extreme length variability, making size-based scheduling difficult in practice. Recent LLM schedulers approximate SJF/SRPT using predicted decode lengths or ra…