1 paper
Rohit Dutta, Paramita Koley, Soham Poddar +5
Speculative decoding has emerged as an effective method to reduce latency and inference cost of LLM inferences. However, there has been inadequate attention towards the energy requ…