1 paper
Weijian Liu, Mingzhen Li, Rui Kang +3
Large language model (LLM) training shall adapt to dynamic resources in shared clusters to tackle the elasticity, including passive preemption and optimistic scaling. State migrati…