1 paper
Daiyaan Arfeen, Dheevatsa Mudigere, Ankit More +3
LLM training is scaled up to 10Ks of GPUs by a mix of data-(DP) and model-parallel (MP) execution. Critical to achieving efficiency is tensor-parallel (TP; a form of MP) execution…