1 paper · 1 filter
Qunyou Liu, Darong Huang, Marina Zapater +1
Large Language Models (LLMs) are becoming the backbone of modern cloud services, yet their inference costs are dominated by GPU energy. Unlike traditional GPU workloads, LLM infere…