2 papers
cs.DC2024
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
Youshao Xiao, Lin Ju, Zhenglei Zhou +8
Many distributed training techniques like Parameter Server and AllReduce have been proposed to take advantage of the increasingly large data and rich features. However, stragglers…
cs.LG2023
Rethinking Memory and Communication Cost for Efficient Large Language Model Training
Chan Wu, Hanxiao Zhang, Lin Ju +8
Recently, various distributed strategies for large language model training have been proposed. However, these methods provided limited solutions for the trade-off between memory co…