20 citations · 20 across the 1 of their papers we have counts for
3 papers
RackSched: A Microsecond-Scale Scheduler for Rack-Scale Computers (Technical Report)
Hang Zhu, Kostis Kaffes, Zixu Chen +4
Low-latency online services have strict Service Level Objectives (SLOs) that require datacenter systems to support high throughput at microsecond-scale tail latency. Dataplane oper…
Blink: Fast and Generic Collectives for Distributed ML
Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee +3
Model parameter synchronization across GPUs introduces high overheads for data-parallel training at scale. Existing parameter synchronization protocols cannot effectively leverage…
numpywren: serverless linear algebra
Vaishaal Shankar, Karl Krauth, Qifan Pu +5
Linear algebra operations are widely used in scientific computing and machine learning applications. However, it is challenging for scientists and data analysts to run linear algeb…