Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
TrainMover: An Interruption-Resilient Runtime for ML Training
ChonLam Lao, Jiaqi Gao, Jiamin Cao +13
Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime r…
cs.DC2025
Decouple and Decompose: Scaling Resource Allocation with DeDe
Zhiying Xu, Minlan Yu, Francis Y. Yan
Efficient resource allocation is essential in cloud systems to facilitate resource sharing among tenants. However, the growing scale of these optimization problems have outpaced co…