3 citations · 3 across the 4 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2024
TrainMover: An Interruption-Resilient Runtime for ML Training
ChonLam Lao, Jiaqi Gao, Jiamin Cao +13
Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime r…
cs.DC2024
Decouple and Decompose: Scaling Resource Allocation with DeDe
Zhiying Xu, Minlan Yu, Francis Y. Yan
Efficient resource allocation is essential in cloud systems to facilitate resource sharing among tenants. However, the growing scale of these optimization problems have outpaced co…