10 citations · 57 across the 26 of their papers we have counts for
8 papers · 1 filter
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
Jiahuan Yu, Mingtao Hu, Zichao Lin +1
Large Language Model (LLM) serving faces a fundamental tension between stringent latency Service Level Objectives (SLOs) and limited GPU memory capacity. When high request rates ex…
VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing
Jiahuan Yu, Aryan Taneja, Junfeng Lin +1
The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although modern serving architectures expose distinct…
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko +4
Deep neural network (DNN) training continues to scale rapidly in terms of model size, data volume, and sequence length, to the point where multiple machines are required to fit lar…
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
Jiangfei Duan, Ziang Song, Xupeng Miao +5
Deep neural networks (DNNs) are becoming progressively large and costly to train. This paper aims to reduce DNN training costs by leveraging preemptible instances on modern clouds,…
Computing in the Era of Large Generative Models: From Cloud-Native to AI-Native
Yao Lu, Song Bian, Lequn Chen +19
In this paper, we investigate the intersection of large generative AI models and cloud-native computing architectures. Recent large models such as ChatGPT, while revolutionary in t…
FedHC: A Scalable Federated Learning Framework for Heterogeneous and Resource-Constrained Clients
Min Zhang, Fuxun Yu, Yongbo Yu +3
Federated Learning (FL) is a distributed learning paradigm that empowers edge devices to collaboratively learn a global model leveraging local data. Simulating FL on GPU is essenti…