36 citations · 46 across the 5 of their papers we have counts for
5 papers
Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
Insu Jang, Zhenning Yang, Zhen Zhang +2
Oobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set…
Memory Disaggregation: Advances and Open Challenges
Hasan Al Maruf, Mosharaf Chowdhury
Compute and memory are tightly coupled within each server in traditional datacenters. Large-scale datacenter operators have identified this coupling as a root cause behind fleet-wi…
Chasing Low-Carbon Electricity for Practical and Sustainable DNN Training
Zhenning Yang, Luoxi Meng, Jae-Won Chung +1
Deep learning has experienced significant growth in recent years, resulting in increased energy consumption and carbon emission from the use of GPUs for training deep neural networ…
FLINT: A Platform for Federated Learning Integration
Ewen Wang, Ajay Kannan, Yuefeng Liang +2
Cross-device federated learning (FL) has been well-studied from algorithmic, system scalability, and training speed perspectives. Nonetheless, moving from centralized training to c…
Orloj: Predictably Serving Unpredictable DNNs
Peifeng Yu, Yuqing Qiu, Xin Jin +1
Existing DNN serving solutions can provide tight latency SLOs while maintaining high throughput via careful scheduling of incoming requests, whose execution times are assumed to be…