75 citations · 116 across the 4 of their papers we have counts for
4 papers
Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review
Jing Su, Chufeng Jiang, Xin Jin +7
This systematic literature review comprehensively examines the application of Large Language Models (LLMs) in forecasting and anomaly detection, highlighting the current state of r…
Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
Insu Jang, Zhenning Yang, Zhen Zhang +2
Oobleck enables resilient distributed training of large DNN models with guaranteed fault tolerance. It takes a planning-execution co-design approach, where it first generates a set…
Energy-Efficient GPU Clusters Scheduling for Deep Learning
Diandian Gu, Xintong Xie, Gang Huang +2
Training deep neural networks (DNNs) is a major workload in datacenters today, resulting in a tremendously fast growth of energy consumption. It is important to reduce the energy c…
MuxFlow: Efficient and Safe GPU Sharing in Large-Scale Production Deep Learning Clusters
Yihao Zhao, Xin Liu, Shufan Liu +5
Large-scale GPU clusters are widely-used to speed up both latency-critical (online) and best-effort (offline) deep learning (DL) workloads. However, most DL clusters either dedicat…