10 citations · 10 across the 1 of their papers we have counts for
2 papers
cs.DC2024
DisDP: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel Training
Mo Sun, Zihan Yang, Changyue Liao +5
Model-sharded data parallelism (MSDP), e.g., ZeRO, evenly shards the model states across all GPUs, and thus has been widely adopted by LLM pre-training, such as Llama and DeepSeek,…
cs.LG2022★ 10 cited
Walle: An End-to-End, General-Purpose, and Large-Scale Production System for Device-Cloud Collaborative Machine Learning
Chengfei Lv, Chaoyue Niu, Renjie Gu +17
To break the bottlenecks of mainstream cloud-based machine learning (ML) paradigm, we adopt device-cloud collaborative ML and build the first end-to-end and general-purpose system,…