2 papers
cs.DC2025
What happens when nanochat meets DiLoCo?
Alexander Acker, Soeren Becker, Sasho Nedelkoski +3
Although LLM training is typically centralized with high-bandwidth interconnects and large compute budgets, emerging methods target communication-constrained training in distribute…
cs.LG2025
Distributed Low-Communication Training with Decoupled Momentum Optimization
Sasho Nedelkoski, Alexander Acker, Odej Kao +2
The training of large models demands substantial computational resources, typically available only in data centers with high-bandwidth interconnects. However, reducing the reliance…