1 paper
Joshua Horswill, Ross Hunter, Matt Clifford +1
We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) toke…