1 paper · 1 filter
Runsheng Benson Guo, Utkarsh Anand, Arthur Chen +1
Training transformer models requires substantial GPU compute and memory resources. In homogeneous clusters, distributed strategies allocate resources evenly, but this approach is i…