3 papers
cs.DC2026
GPU Memory and Utilization Estimation for Training-Aware Resource Management: Opportunities and Limitations
Ehsan Yousefzadeh-Asl-Miandoab, Reza Karimzadeh, Danyal Yorulmaz +2
Collocating deep learning training tasks improves GPU utilization but risks resource contention, severe slowdowns, and out-of-memory (OOM) failures. Accurate memory estimation is e…
cs.DC2025
CARMA: Collocation-Aware Resource Manager
Ehsan Yousefzadeh-Asl-Miandoab, Florina M. Ciorba, Pınar Tözün
GPUs running deep learning (DL) workloads are frequently underutilized. Collocating multiple DL training tasks on the same GPU can improve utilization but introduces two key risks:…
cs.LG2022
An Analysis of Collocation on GPUs for Deep Learning Training
Ties Robroek, Ehsan Yousefzadeh-Asl-Miandoab, Pınar Tözün
Deep learning training is an expensive process that extensively uses GPUs, but not all model training saturates modern powerful GPUs. Multi-Instance GPU (MIG) is a new technology i…