9 citations · 9 across the 3 of their papers we have counts for
1 paper · 1 filter
Daiyaan Arfeen, Dheevatsa Mudigere, Ankit More +3
LLM training is scaled up to 10Ks of GPUs by a mix of data-(DP) and model-parallel (MP) execution. Critical to achieving efficiency is tensor-parallel (TP; a form of MP) execution…