2 papers
cs.LG2026
Communication-Efficient LLM Adaptation over Decentralized GPU Meshes
Sameera Ramasinghe, Shamane Siriwardhana, Thalaiyasingam Ajanthan +8
Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes…
cs.LG2026
Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language Models
Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi +9
Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homo…