19 citations · 55 across the 6 of their papers we have counts for
8 papers
RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers (Technical Report)
Hang Zhu, Kostis Kaffes, Zixu Chen +4
Low-latency online services have strict Service Level Objectives (SLOs) that require datacenter systems to support high throughput at microsecond-scale tail latency. Dataplane oper…
On Efficient Constructions of Checkpoints
Yu Chen, Zhenming Liu, Bin Ren +1
Efficient construction of checkpoints/snapshots is a critical tool for training and diagnosing deep learning models. In this paper, we propose a lossy compression scheme for checkp…
Is Network the Bottleneck of Distributed Training?
Zhen Zhang, Chaokun Chang, Haibin Lin +3
Recently there has been a surge of research on improving the communication efficiency of distributed training. However, little work has been done to systematically understand wheth…
Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict Detection
Hang Zhu, Zhihao Bai, Jialin Li +4
Distributed storage employs replication to mask failures and improve availability. However, these systems typically exhibit a hard tradeoff between consistency and performance. Ens…
Neural Packet Classification
Eric Liang, Hang Zhu, Xin Jin +1
Packet classification is a fundamental problem in computer networking. This problem exposes a hard tradeoff between the computation and state complexity, which makes it particularl…
DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching
Zaoxing Liu, Zhihao Bai, Zhenming Liu +5
Load balancing is critical for distributed storage to meet strict service-level objectives (SLOs). It has been shown that a fast cache can guarantee load balancing for a clustered…