4 papers
Benchmarking the Energy Savings with Speculative Decoding Strategies
Rohit Dutta, Paramita Koley, Soham Poddar +5
Speculative decoding has emerged as an effective method to reduce latency and inference cost of LLM inferences. However, there has been inadequate attention towards the energy requ…
RSTGCN: Railway-centric Spatio-Temporal Graph Convolutional Network for Train Delay Prediction
Koyena Chowdhury, Paramita Koley, Abhijnan Chakraborty +1
Accurate prediction of train delays is critical for efficient railway operations. While earlier approaches have largely focused on forecasting the exact delays of individual trains…
Brevity is the soul of sustainability: Characterizing LLM response lengths
Soham Poddar, Paramita Koley, Janardan Misra +4
A significant portion of the energy consumed by Large Language Models (LLMs) arises from their inference processes; hence developing energy-efficient methods for inference is cruci…
Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models
Soham Poddar, Paramita Koley, Janardan Misra +3
Large language models (LLMs) are increasingly recognized for their exceptional generative capabilities and versatility across various tasks. However, the high inference costs assoc…