3 papers
cs.LG2026
Understanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
Julien Delavande, Regis Pierrard, Sasha Luccioni
Large Language Models (LLMs) are increasingly deployed in production, contributing towards shifting the burden in terms of computational resources and energy demands from training…
cs.LG2026
Small Talk, Big Impact: The Energy Cost of Thanking AI
Julien Delavande, Regis Pierrard, Sasha Luccioni
Being polite is free - or is it? In this paper, we quantify the energy cost of seemingly innocuous messages such as ``thank you'' when interacting with large language models, often…
cs.LG2025
Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models
Julien Delavande, Regis Pierrard, Sasha Luccioni
Recent advances in text-to-video (T2V) generation have enabled the creation of high-fidelity, temporally coherent clips from natural language prompts. Yet these systems come with s…