19 citations · 55 across the 6 of their papers we have counts for
5 papers · 1 filter
RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers (Technical Report)
Hang Zhu, Kostis Kaffes, Zixu Chen +4
Low-latency online services have strict Service Level Objectives (SLOs) that require datacenter systems to support high throughput at microsecond-scale tail latency. Dataplane oper…
Is Network the Bottleneck of Distributed Training?
Zhen Zhang, Chaokun Chang, Haibin Lin +3
Recently there has been a surge of research on improving the communication efficiency of distributed training. However, little work has been done to systematically understand wheth…
Harmonia: Near-Linear Scalability for Replicated Storage with In-Network Conflict Detection
Hang Zhu, Zhihao Bai, Jialin Li +4
Distributed storage employs replication to mask failures and improve availability. However, these systems typically exhibit a hard tradeoff between consistency and performance. Ens…
DistCache: Provable Load Balancing for Large-Scale Storage Systems with Distributed Caching
Zaoxing Liu, Zhihao Bai, Zhenming Liu +5
Load balancing is critical for distributed storage to meet strict service-level objectives (SLOs). It has been shown that a fast cache can guarantee load balancing for a clustered…
NetChain: Scale-Free Sub-RTT Coordination (Extended Version)
Xin Jin, Xiaozhou Li, Haoyu Zhang +5
Coordination services are a fundamental building block of modern cloud systems, providing critical functionalities like configuration management and distributed locking. The major…