2 papers
cs.LG2025
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
Ling Team, Binwei Zeng, Chao Huang +71
In this technical report, we tackle the challenges of training large-scale Mixture of Experts (MoE) models, focusing on overcoming cost inefficiency and resource limitations preval…
cs.LG2024
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
Youshao Xiao, Zhenglei Zhou, Fagui Mao +6
Recently, ChatGPT or InstructGPT like large language models (LLM) has made a significant impact in the AI world. Many works have attempted to reproduce the complex InstructGPT's tr…