1 paper · 1 filter
Xiang Liu, Shimiao Yuan, Zhenheng Tang +5
LLM inference is still evaluated mainly as a model or software problem: accuracy, latency, throughput, and hardware utilization. This is incomplete. At deployment scale, the releva…