12 citations · 12 across the 2 of their papers we have counts for
3 papers
cs.DC2024
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
Youshao Xiao, Lin Ju, Zhenglei Zhou +8
Many distributed training techniques like Parameter Server and AllReduce have been proposed to take advantage of the increasingly large data and rich features. However, stragglers…
cs.LG2024★ 12 cited
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
Youshao Xiao, Shangchun Zhao, Zhenglei Zhou +5
Recently, a new paradigm, meta learning, has been widely applied to Deep Learning Recommendation Models (DLRM) and significantly improves statistical performance, especially in col…
cs.LG2023
An Adaptive Placement and Parallelism Framework for Accelerating RLHF Training
Youshao Xiao, Zhenglei Zhou, Fagui Mao +6
Recently, ChatGPT or InstructGPT like large language models (LLM) has made a significant impact in the AI world. Many works have attempted to reproduce the complex InstructGPT's tr…