17 citations · 21 across the 2 of their papers we have counts for
2 papers
cs.NI2023★ 4 cited
Rail-only: A Low-Cost High-Performance Network for Training LLMs with Trillion Parameters
Weiyang Wang, Manya Ghobadi, Kayvon Shakeri +2
This paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a nov…
cs.NI2022★ 17 cited
TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
Weiyang Wang, Moein Khazraee, Zhizhen Zhong +5
We propose TopoOpt, a novel direct-connect fabric for deep neural network (DNN) training workloads. TopoOpt co-optimizes the distributed training process across three dimensions: c…