1 paper
Xiaosong Chen, Shaoheng Nie, Zhongmin Zhao +6
With the rapid advancement of accelerator technologies, pre-training large language models (LLMs) on heterogeneous accelerator clusters has become increasingly crucial for maximizi…