4 citations · 4 across the 4 of their papers we have counts for
4 papers
WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-Martínez
Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to…
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-Martínez
While the large energy consumption of Large Language Models (LLMs) is recognized by the community, system operators lack guidance for energy-efficient LLM inference deployments tha…
Bench360: Benchmarking Local LLM Inference from 360 Degrees
Linus Stuhlmann, Mauricio Fadel Argerich, Jonathan Fürst
Running LLMs locally has become increasingly common, but users face a complex design space across models, quantization levels, inference engines, and serving scenarios. Existing in…
Reinforcement Learning Based Orchestration for Elastic Services
M. Fadel Argerich, B. Cheng, J. Fürst
Due to the highly variable execution context in which edge services run, adapting their behavior to the execution context is crucial to comply with their requirements. However, ada…