3 papers
cs.LG2025
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
Zicong Ye, Kunming Zhang, Guoming Tang
The explosive growth of interactive Large Language Models (LLMs) has placed unprecedented demands for low latency on cloud GPUs, forcing them into high-power modes and causing esca…
cs.DC2025
GPUnion: Autonomous GPU Sharing on Campus
Yufang Li, Yuanbo Zhang, Hanlong Liao +2
A pronounced imbalance in GPU resources exists on campus, where some laboratories own underutilized servers while others lack the compute needed for AI research. GPU sharing can al…
cs.DC2025
AIMeter: Measuring, Analyzing, and Visualizing Energy and Carbon Footprint of AI Workloads
Hongzhen Huang, Kunming Zhang, Hanlong Liao +2
The rapid advancement of AI, particularly large language models (LLMs), has raised significant concerns about the energy use and carbon emissions associated with model training and…