3 papers
cs.LG2024
DiPaCo: Distributed Path Composition
Arthur Douillard, Qixuan Feng, Andrei A. Rusu +7
Progress in machine learning (ML) has been fueled by scaling neural network models. This scaling has been enabled by ever more heroic feats of engineering, necessary for accommodat…
cs.LG2024
Asynchronous Local-SGD Training for Language Modeling
Bo Liu, Rachita Chhaparia, Arthur Douillard +5
Local stochastic gradient descent (Local-SGD), also referred to as federated averaging, is an approach to distributed optimization where each device performs more than one SGD upda…
cs.LG2023
DiLoCo: Distributed Low-Communication Training of Language Models
Arthur Douillard, Qixuan Feng, Andrei A. Rusu +6
Large language models (LLM) have become a critical component in many applications of machine learning. However, standard approaches to training LLM require a large number of tightl…