14 citations · 60 across the 17 of their papers we have counts for
1 paper · 1 filter
Xiang Liu, Shimiao Yuan, Zhenheng Tang +5
LLM inference is still evaluated mainly as a model or software problem: accuracy, latency, throughput, and hardware utilization. This is incomplete. At deployment scale, the releva…