Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
FCDP: Fully Cached Data Parallel for Communication-Avoiding Large-Scale Training
Gyeongseo Park, Eungyeong Lee, Song-woo Sok +5
Training billion-parameter models requires distributing model states across GPUs using fully sharded data parallel (i.e., ZeRO-3). While ZeRO-3 succeeds on clusters with high-bandw…
cs.DC2024
DeepVM: Integrating Spot and On-Demand VMs for Cost-Efficient Deep Learning Clusters in the Cloud
Yoochan Kim, Kihyun Kim, Yonghyeon Cho +7
Distributed Deep Learning (DDL), as a paradigm, dictates the use of GPU-based clusters as the optimal infrastructure for training large-scale Deep Neural Networks (DNNs). However,…