13 papers
On Topology's Role in ML Training Performance
Sarah McClure, Tegan Wilson, Brad Karp +4
Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basi…
EnCoR: An end-to-end architecture for simplifying cellular networks
Wesley Woo, Zhuowei Wen, Monniiesh Velmurugan +4
Since their creation, cellular networks have made in-network mobility support a key feature of their service model. While this approach provides seamless connectivity for legacy tr…
Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving
Junsun Choi, Sam Son, Sunjin Choi +5
Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of LLM serving runtime. This has…
Load Balancing for AI Training Workloads
Sarah McClure, Evyatar Cohen, Alex Shpiner +4
The extreme bandwidth demands of AI training has made load-balancing a critical component in AI fabrics, and a variety of load-balancing designs have emerged in recent work from bo…
TURBO: Utility-Aware Bandwidth Allocation for Cloud-Augmented Autonomous Control
Peter Schafhalter, Alexander Krentsel, Hongbo Wei +4
Autonomous driving system progress has been driven by improvements in machine learning models, whose computational demands now exceed what edge devices alone can provide. The cloud…
Delta Fair Sharing: Performance Isolation for Multi-Tenant Storage Systems
Tyler Griggs, Soujanya Ponnapalli, Dev Bali +8
Modern storage systems, often deployed to support multiple tenants in the cloud, must provide performance isolation. Unfortunately, traditional approaches such as fair sharing do n…