1 paper · 1 filter
Jaeyong Song, Jinkyu Yim, Jaewon Jung +4
In training of modern large natural language processing (NLP) models, it has become a common practice to split models using 3D parallelism to multiple GPUs. Such technique, however…