Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
MuLoCo: Muon is a practical inner optimizer for DiLoCo
Benjamin Thérien, Xiaolong Huang, Aaron Defazio +2
DiLoCo is a powerful framework for training large language models (LLMs), enabling larger optimal batch sizes and increased accelerator utilization under networking constraints. Ho…
cs.LG2026
Model Parallelism With Subnetwork Data Parallelism
Vaibhav Singh, Zafir Khalid, Pietro Cagnasso +2
Pre-training large neural networks at scale imposes heavy memory demands on accelerators and often requires costly communication. We introduce Subnetwork Data Parallelism (SDP), a…