1 citations · 1 across the 4 of their papers we have counts for
3 papers · 1 filter
PCCL: Process Group-Aware Scalable and Generic Collective Algorithm Synthesizer
William Won, Kartik Lakhotia, Madhu Kumar +2
Distributed machine learning has become increasingly important due to the massive scale of large-scale generative models. Both model parameters and data are distributed across many…
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
William Won, Midhilesh Elavazhagan, Sudarshan Srinivasan +2
The surge of artificial intelligence, particularly large language models, has driven the rapid development of large-scale machine learning clusters. Executing distributed models on…
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
William Won, Saeed Rashidi, Sudarshan Srinivasan +1
As model sizes in machine learning continue to scale, distributed training is necessary to accommodate model weights within each device and to reduce training time. However, this c…