4 papers
BandPilot: Toward Performance- and Contention-Aware GPU Dispatching in AI Clusters
Kunming Zhang, Hanlong Liao, Junyu Xue +2
Modern multi-tenant AI clusters are increasingly communication-bound, driven by high-volume and multi-round GPU-to-GPU collective communication. Consequently, the GPU dispatcher's…
GPUnion: Autonomous GPU Sharing on Campus
Yufang Li, Yuanbo Zhang, Hanlong Liao +2
A pronounced imbalance in GPU resources exists on campus, where some laboratories own underutilized servers while others lack the compute needed for AI research. GPU sharing can al…
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
Hongzhen Huang, Kunming Zhang, Hanlong Liao +2
The rapid advancement of AI, particularly large language models (LLMs), has raised significant concerns about the energy use and carbon emissions associated with model training and…
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
Zicong Ye, Kunming Zhang, Guoming Tang
The explosive growth of interactive Large Language Models (LLMs) has placed unprecedented demands for low latency on cloud GPUs, forcing them into high-power modes and causing esca…