Publications (12)
My CXL Pool Obviates Your PCIe Switch
Yuhong Zhong, Daniel S. Berger, Pantea Zardoshti +5
Pooling PCIe devices across multiple hosts offers a promising solution to mitigate stranded I/O resources, enhance device utilization, address device failures, and reduce total cos…
Unlocking the Power of Inline Floating-Point Operations on Programmable Switches
Yifan Yuan, Omar Alama, Amedeo Sapio +5
The advent of switches with programmable dataplanes has enabled the rapid development of new network functionality, as well as providing a platform for acceleration of a broad rang…
ORCA: A Network and Architecture Co-design for Offloading us-scale Datacenter Applications
Yifan Yuan, Jinghan Huang, Yan Sun +7
Responding to the "datacenter tax" and "killer microseconds" problems for datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless,…
Cloud Collectives: Towards Cloud-aware Collectives forML Workloads with Rank Reordering
Liang Luo, Jacob Nelson, Arvind Krishnamurthy +1
ML workloads are becoming increasingly popular in the cloud. Good cloud training performance is contingent on efficient parameter exchange among VMs. We find that Collectives, the…
TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches
Aashaka Shah, Vijay Chidambaram, Meghan Cowan +6
Machine learning models are increasingly being trained across multiple GPUs and servers. In this setting, data is transferred between GPUs using communication collectives such as A…
Scaling Distributed Machine Learning with In-Network Aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho +7
Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a p…
Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training
Liang Luo, Jacob Nelson, Luis Ceze +2
Distributed deep neural network (DDNN) training constitutes an increasingly important workload that frequently runs in the cloud. Larger DNN models and faster compute engines are s…
Parameter Box: High Performance Parameter Servers for Efficient Distributed Deep Neural Network Training
Liang Luo, Jacob Nelson, Luis Ceze +2
Most work in the deep learning systems community has focused on faster inference, but arriving at a trained model requires lengthy experiments. Accelerating training lets developer…
Beehive: A Flexible Network Stack for Direct-Attached Accelerators
Katie Lim, Matthew Giordano, Theano Stavrinos +4
Direct-attached accelerators, where application accelerators are directly connected to the datacenter network via a hardware network stack, offer substantial benefits in terms of r…
Synthesizing Optimal Collective Algorithms
Zixian Cai, Zhengyang Liu, Saeed Maleki +4
Collective communication algorithms are an important component of distributed computation. Indeed, in the case of deep-learning, collective communication is the Amdahl's bottleneck…
Bundled References: An Abstraction for Highly-Concurrent Linearizable Range Queries
Jacob Nelson, Ahmed Hassan, Roberto Palmieri
We present bundled references, a new building block to provide linearizable range query operations for highly concurrent linked data structures. Bundled references allow range quer…
Hybrid Computing for Interactive Datacenter Applications
Pratyush Patel, Katie Lim, Kushal Jhunjhunwalla +5
Field-Programmable Gate Arrays (FPGAs) are more energy efficient and cost effective than CPUs for a wide variety of datacenter applications. Yet, for latency-sensitive and bursty w…