1 paper · 1 filter
Mingkai Zheng, Junlin Chen, Haotian Xie +1
Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronizatio…