2 papers
cs.DC2026
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
Ehsan Yousefzadeh-Asl-Miandoab, Reza Karimzadeh, Danyal Yorulmaz +2
Collocating deep learning training tasks improves GPU utilization but risks resource contention, severe slowdowns, and out-of-memory (OOM) failures. Accurate memory estimation is e…
cs.DC2026
CARMA: Collocation-Aware Resource Manager
Ehsan Yousefzadeh-Asl-Miandoab, Florina M. Ciorba, Pınar Tözün
GPUs running deep learning (DL) workloads are frequently underutilized. Collocating multiple DL training tasks on the same GPU can improve utilization but introduces two key risks:…