1 paper
Weiyang Wang, Manya Ghobadi, Kayvon Shakeri +2
This paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a nov…