1 paper · 1 filter
Houming Wu, Ling Chen
Training large language models (LLMs) is fundamentally constrained by limited device memory and costly inter-device communication. Although pipeline parallelism alleviates memory p…