papers

Publications (12)

cs.OS2025

My CXL Pool Obviates Your PCIe Switch

Yuhong Zhong, Daniel S. Berger, Pantea Zardoshti +5

Pooling PCIe devices across multiple hosts offers a promising solution to mitigate stranded I/O resources, enhance device utilization, address device failures, and reduce total cos…

cs.NI2021

Unlocking the Power of Inline Floating-Point Operations on Programmable Switches

Yifan Yuan, Omar Alama, Amedeo Sapio +5

The advent of switches with programmable dataplanes has enabled the rapid development of new network functionality, as well as providing a platform for acceleration of a broad rang…

cs.AR2022

ORCA: A Network and Architecture Co-design for Offloading us-scale Datacenter Applications

Yifan Yuan, Jinghan Huang, Yan Sun +7

Responding to the "datacenter tax" and "killer microseconds" problems for datacenter applications, diverse solutions including Smart NIC-based ones have been proposed. Nonetheless,…

cs.DC2021

Cloud Collectives: Towards Cloud-aware Collectives forML Workloads with Rank Reordering

Liang Luo, Jacob Nelson, Arvind Krishnamurthy +1

ML workloads are becoming increasingly popular in the cloud. Good cloud training performance is contingent on efficient parameter exchange among VMs. We find that Collectives, the…

cs.DC2022

TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches

Aashaka Shah, Vijay Chidambaram, Meghan Cowan +6

Machine learning models are increasingly being trained across multiple GPUs and servers. In this setting, data is transferred between GPUs using communication collectives such as A…

cs.DC2020

Scaling Distributed Machine Learning with In-Network Aggregation

Amedeo Sapio, Marco Canini, Chen-Yu Ho +7

Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a p…

cs.DC2020

Parameter Hub: a Rack-Scale Parameter Server for Distributed Deep Neural Network Training

Liang Luo, Jacob Nelson, Luis Ceze +2

Distributed deep neural network (DDNN) training constitutes an increasingly important workload that frequently runs in the cloud. Larger DNN models and faster compute engines are s…

cs.DC2020

Parameter Box: High Performance Parameter Servers for Efficient Distributed Deep Neural Network Training

Liang Luo, Jacob Nelson, Luis Ceze +2

Most work in the deep learning systems community has focused on faster inference, but arriving at a trained model requires lengthy experiments. Accelerating training lets developer…

cs.AR2024

Beehive: A Flexible Network Stack for Direct-Attached Accelerators

Katie Lim, Matthew Giordano, Theano Stavrinos +4

Direct-attached accelerators, where application accelerators are directly connected to the datacenter network via a hardware network stack, offer substantial benefits in terms of r…

cs.DC2021

Synthesizing Optimal Collective Algorithms

Zixian Cai, Zhengyang Liu, Saeed Maleki +4

Collective communication algorithms are an important component of distributed computation. Indeed, in the case of deep-learning, collective communication is the Amdahl's bottleneck…

cs.DS2020

Bundled References: An Abstraction for Highly-Concurrent Linearizable Range Queries

Jacob Nelson, Ahmed Hassan, Roberto Palmieri

We present bundled references, a new building block to provide linearizable range query operations for highly concurrent linked data structures. Bundled references allow range quer…

cs.DC2023

Hybrid Computing for Interactive Datacenter Applications

Pratyush Patel, Katie Lim, Kushal Jhunjhunwalla +5

Field-Programmable Gate Arrays (FPGAs) are more energy efficient and cost effective than CPUs for a wide variety of datacenter applications. Yet, for latency-sensitive and bursty w…