3 papers
cs.DC2026
HeLoCo: Efficient asynchronous low-communication training under data and device heterogeneity
Abdullah Al Asif, Patrick Diem, Juan Pablo Muñoz +3
Distributed Low-Communication (DiLoCo) training reduces communication overhead by allowing workers to perform multiple local optimization steps before sending pseudo-gradients to a…
cs.DC2026
SuperSFL: Resource-Heterogeneous Federated Split Learning with Weight-Sharing Super-Networks
Abdullah Al Asif, Sixing Yu, Juan Pablo Munoz +2
SplitFed Learning (SFL) combines federated learning and split learning to enable collaborative training across distributed edge devices; however, it faces significant challenges in…
cs.CL2024
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
Branden Butler, Sixing Yu, Arya Mazaheri +1
Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from C…