activity
20242026
collaborators

6 papers

cs.CL2026

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model

Mohammed Sabry, Sean Augenstein, Keith Rush +1

We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a coherent larger model. We introduce Mixt…

cs.CL2026

Decoupled DiLoCo for Resilient Distributed Pre-training

Arthur Douillard, Keith Rush, Yani Donchev +14

Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across accelerators. Due to this co…

cs.LG2025

Correlated Noise Mechanisms for Differentially Private Learning

Krishna Pillutla, Jalaj Upadhyay, Christopher A. Choquette-Choo +9

This monograph explores the design and analysis of correlated noise mechanisms for differential privacy (DP), focusing on their application to private training of AI and machine le…

cs.LG2025

Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo

Zachary Charles, Gabriel Teston, Lucio Dery +5

As we scale to more massive machine learning models, the frequent synchronization demands inherent in data-parallel approaches create significant slowdowns, posing a critical chall…

cs.CL2025

Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch

Arthur Douillard, Yanislav Donchev, Keith Rush +11

Training of large language models (LLMs) is typically distributed across a large number of accelerators to reduce training time. Since internal states and parameter gradients need…

cs.LG2024

Federated Automatic Differentiation

Keith Rush, Zachary Charles, Zachary Garrett

Federated learning (FL) is a general framework for learning across an axis of group partitioned data (heterogeneous clients) while preserving data privacy, under the orchestration…