most citedOobleck: Resilient Distributed Training of Large Models Using Pipeline Templates

36 citations · 46 across the 5 of their papers we have counts for

collaborators

5 papers

cs.DC202336 cited

Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates

Insu Jang, Zhenning Yang, Zhen Zhang +2

Oobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set…

cs.DC20231 cited

Memory Disaggregation: Advances and Open Challenges

Hasan Al Maruf, Mosharaf Chowdhury

Compute and memory are tightly coupled within each server in traditional datacenters. Large-scale datacenter operators have identified this coupling as a root cause behind fleet-wi…

cs.LG20233 cited

Chasing Low-Carbon Electricity for Practical and Sustainable DNN Training

Zhenning Yang, Luoxi Meng, Jae-Won Chung +1

Deep learning has experienced significant growth in recent years, resulting in increased energy consumption and carbon emission from the use of GPUs for training deep neural networ…

cs.LG20236 cited

FLINT: A Platform for Federated Learning Integration

Ewen Wang, Ajay Kannan, Yuefeng Liang +2

Cross-device federated learning (FL) has been well-studied from algorithmic, system scalability, and training speed perspectives. Nonetheless, moving from centralized training to c…

cs.DC2022

Orloj: Predictably Serving Unpredictable DNNs

Peifeng Yu, Yuqing Qiu, Xin Jin +1

Existing DNN serving solutions can provide tight latency SLOs while maintaining high throughput via careful scheduling of incoming requests, whose execution times are assumed to be…