2 papers
cs.DC2025
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
Runsheng Benson Guo, Utkarsh Anand, Khuzaima Daudjee +1
Large language models (LLMs) require vast amounts of GPU compute to train, but limited availability and high costs of GPUs make homogeneous clusters impractical for many organizati…
cs.DC2024
Cephalo: Harnessing Heterogeneous GPU Clusters for Training Transformer Models
Runsheng Benson Guo, Utkarsh Anand, Arthur Chen +1
Training transformer models requires substantial GPU compute and memory resources. In homogeneous clusters, distributed strategies allocate resources evenly, but this approach is i…