1 paper
Weifang Hu, Xuanhua Shi, Yunkai Zhang +7
Optimizing the parallel training of large models requires exploring intra-operator parallelism plans for a computation graph that typically contains tens of thousands of primitive…