19 citations · 19 across the 1 of their papers we have counts for
4 papers
The nanoPU: Redesigning the CPU-Network Interface to Minimize RPC Tail Latency
Stephen Ibanez, Alex Mallery, Serhat Arslan +4
The nanoPU is a new networking-optimized CPU designed to minimize tail latency for RPCs. By bypassing the cache and memory hierarchy, the nanoPU directly places arriving messages i…
Scaling Distributed Machine Learning with In-Network Aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho +7
Training machine learning models in parallel is an increasingly important workload. We accelerate distributed parallel training by designing a communication primitive that uses a p…
DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching
Zaoxing Liu, Zhihao Bai, Zhenming Liu +5
Load balancing is critical for distributed storage to meet strict service-level objectives (SLOs). It has been shown that a fast cache can guarantee load balancing for a clustered…
NetChain: Scale-Free Sub-RTT Coordination (Extended Version)
Xin Jin, Xiaozhou Li, Haoyu Zhang +5
Coordination services are a fundamental building block of modern cloud systems, providing critical functionalities like configuration management and distributed locking. The major…