2 papers
cs.CL2026
Decoupled DiLoCo for Resilient Distributed Pre-training
Arthur Douillard, Keith Rush, Yani Donchev +14
Modern large-scale language model pre-training relies heavily on the single program multiple data (SPMD) paradigm, which requires tight coupling across accelerators. Due to this co…
cs.LG2026
Latent Space Communication via K-V Cache Alignment
Lucio M. Dery, Zohar Yahav, Henry Prior +3
Solving increasingly complex problems with large language models (LLMs) necessitates a move beyond individual models and towards multi-model systems that can effectively collaborat…