1 paper · 1 filter
Guangyu Xiang, Lin Zhang, Haoxuan Yu +3
Overlapping communication and computation operators is a common practice to hide communication overheads, accelerating large language models (LLMs) training on GPU clusters. Existi…