2 citations · 3 across the 3 of their papers we have counts for
3 papers
PRIME: Pseudo-Random Integrated Multi-Part Entropy for Adaptive Packet Spraying in AI/ML Data centers
Ashkan Sobhani, Sogand Sadrhaghighi, Xingjun Chu
Large-scale distributed training in production data centers place significant demands on network infrastructure. In particular, significant load balancing challenges arise when pro…
P/D-Serve: Serving Disaggregated Large Language Model at Scale
Yibo Jin, Tao Wang, Huimin Lin +27
Serving disaggregated large language models (LLMs) over tens of thousands of xPU devices (GPUs or NPUs) with reliable performance faces multiple challenges. 1) Ignoring the diversi…
On the Burstiness of Distributed Machine Learning Traffic
Natchanon Luangsomboon, Fahimeh Fazel, Jörg Liebeherr +3
Traffic from distributed training of machine learning (ML) models makes up a large and growing fraction of the traffic mix in enterprise data centers. While work on distributed ML…