3 papers
cs.LG2025
Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
Jeffrey Amico, Gabriel Passamani Andrade, John Donaghy +12
Post-training language models (LMs) with reinforcement learning (RL) can enhance their complex reasoning capabilities without supervised fine-tuning, as demonstrated by DeepSeek-R1…
cs.LG2025
NoLoCo: No-all-reduce Low Communication Training Method for Large Models
Jari Kolehmainen, Nikolay Blagoev, John Donaghy +2
Training large language models is generally done via optimization methods on clusters containing tens of thousands of accelerators, communicating over a high-bandwidth interconnect…
cs.LG2025
HDEE: Heterogeneous Domain Expert Ensemble
OÄuzhan Ersoy, Jari Kolehmainen, Gabriel Passamani Andrade
Training dense LLMs requires enormous amounts of data and centralized compute, which introduces fundamental bottlenecks and ever-growing costs for large models. Several studies aim…