Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-MartÃnez
Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to…
cs.DC2026
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-MartÃnez
While the large energy consumption of Large Language Models (LLMs) is recognized by the community, system operators lack guidance for energy-efficient LLM inference deployments tha…