24 citations · 25 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 24 cited
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong +29
We present the design, implementation and engineering experience in building and deploying MegaScale, a production system for training large language models (LLMs) at the scale of…
cs.DC2023★ 1 cited
Baechi: Fast Device Placement of Machine Learning Graphs
Beomyeol Jeon, Linda Cai, Chirag Shetty +6
Machine Learning graphs (or models) can be challenging or impossible to train when either devices have limited memory, or models are large. To split the model across devices, learn…