961 citations · 1.1k across the 19 of their papers we have counts for
1 paper · 2 filters
Qiyang Ding, Pengfei Zheng, Shreyas Kudari +2
Accommodating long-running deep learning (DL) training and inference jobs is challenging on GPU clusters that use traditional batch schedulers, such as Slurm. Given fixed wall cloc…