1 paper
Lang Xu, Quentin Anthony, Qinghua Zhou +5
Data Parallelism (DP), Tensor Parallelism (TP), and Pipeline Parallelism (PP) are the three strategies widely adopted to enable fast and efficient Large Language Model (LLM) traini…