3 papers
cs.DC2026
BandPilot: Toward Performance- and Contention-Aware GPU Dispatching in AI Clusters
Kunming Zhang, Hanlong Liao, Junyu Xue +2
Modern multi-tenant AI clusters are increasingly communication-bound, driven by high-volume and multi-round GPU-to-GPU collective communication. Consequently, the GPU dispatcher's…
cs.DC2025
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
Hongzhen Huang, Kunming Zhang, Hanlong Liao +2
The rapid advancement of AI, particularly large language models (LLMs), has raised significant concerns about the energy use and carbon emissions associated with model training and…
cs.LG2025
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
Zicong Ye, Kunming Zhang, Guoming Tang
The explosive growth of interactive Large Language Models (LLMs) has placed unprecedented demands for low latency on cloud GPUs, forcing them into high-power modes and causing esca…