1 paper
Qiaoling Chen, Qinghao Hu, Guoteng Wang +8
Training large language models (LLMs) encounters challenges in GPU memory consumption due to the high memory requirements of model states. The widely used Zero Redundancy Optimizer…