2 papers
cs.LG2026
Benchmarking the Energy Savings with Speculative Decoding Strategies
Rohit Dutta, Paramita Koley, Soham Poddar +5
Speculative decoding has emerged as an effective method to reduce latency and inference cost of LLM inferences. However, there has been inadequate attention towards the energy requ…
cs.CL2025
Brevity is the soul of sustainability: Characterizing LLM response lengths
Soham Poddar, Paramita Koley, Janardan Misra +4
A significant portion of the energy consumed by Large Language Models (LLMs) arises from their inference processes; hence developing energy-efficient methods for inference is cruci…