1 paper · 1 filter
Alicia Golden, Michael Kuchnik, Samuel Hsia +4
Large model training beyond tens of thousands of GPUs is an uncharted territory. At such scales, disruptions to the training process are not a matter of if, but a matter of when --…