Publications (41)
Zero-CPU Collection with Direct Telemetry Access
Jonatan Langlet, Ran Ben Basat, Sivaramakrishnan Ramanathan +4
Programmable switches are driving a massive increase in fine-grained measurements. This puts significant pressure on telemetry collectors that have to process reports from many swi…
On Topology's Role in ML Training Performance
Sarah McClure, Tegan Wilson, Brad Karp +4
Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basi…
TrainMover: An Interruption-Resilient Runtime for ML Training
ChonLam Lao, Jiaqi Gao, Jiamin Cao +13
Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime r…
Federated Learning Clients Clustering with Adaptation to Data Drifts
Minghao Li, Dmitrii Avdiukhin, Rana Shahout +3
Federated Learning (FL) trains deep models across edge devices without centralizing raw data, preserving user privacy. However, client heterogeneity slows down convergence and limi…
SVR-MAD: A Bayesian-Inspired Framework for Posterior-Guided Multi-Agent Debate
Weifan Jiang, Rana Shahout, Minghao Li +4
Multi-Agent Debate (MAD) improves LLM-agent accuracy but suffers from rapid context growth, limiting scalability in larger multi-agent settings. Existing methods prune low-utility…
Collective Communication for 100k+ GPUs
Min Si, Pavan Balaji, Yongzhou Chen +36
The increasing scale of large language models (LLMs) necessitates highly efficient collective communication frameworks, particularly as training workloads extend to hundreds of tho…