1 paper
Danush Khanna, Aditya Kumar Guru, Srivarshinee Sridhar +7
Inference accounts for the majority of latency and energy consumption in large language model (LLM) deployments, often exceeding 90% of total cost. While training-time efficiency h…