1 paper
Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan +8
Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes…