2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.NI2026★ 2 cited
Resilient AI Supercomputer Networking using MRC and SRv6
Joao Araujo, Alex Chow, Mark Handley +47
Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport proto…
cs.DB2024
Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service
Deepak Ravikumar, Alex Yeo, Yiwen Zhu +11
The proliferation of big data and analytic workloads has driven the need for cloud compute and cluster-based job processing. With Apache Spark, users can process terabytes of data…