4 papers
Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
Joel Lidin, Amir Sarfi, Erfan Miahi +6
Recently, there has been increased interest in globally distributed training, which has the promise to both reduce training costs and democratize participation in building large-sc…
Heterogeneous Low-Bandwidth Pre-Training of LLMs
Yazan Obeidi, Amir Sarfi, Joel Lidin +2
Pre-training large language models (LLMs) increasingly requires distributed compute, yet bandwidth constraints make it difficult to scale beyond well-provisioned datacenters-especi…
Overcoming the Communication-Performance Tradeoff in LLM Pretraining
Amir Sarfi, Benjamin Thérien, Joel Lidin +1
Communication-efficient distributed training algorithms (e.g., DiLoCo) have received considerable interest due to their benefits for training large language models (LLMs) in bandwi…
Incentivizing Permissionless Distributed Learning of LLMs
Joel Lidin, Amir Sarfi, Evangelos Pappas +3
We describe an incentive system for distributed deep learning of foundational models where peers are rewarded for contributions. The incentive system, \textit{Gauntlet}, has been d…