12 citations · 23 across the 44 of their papers we have counts for
Showing cs.CEShow all
2 papers · 1 filter
cs.CE2026
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
Xiang Liu, Shimiao Yuan, Zhenheng Tang +5
LLM inference is still evaluated mainly as a model or software problem: accuracy, latency, throughput, and hardware utilization. This is incomplete. At deployment scale, the releva…
cs.CE2024
Task Scheduling for Efficient Inference of Large Language Models on Single Moderate GPU Systems
Wenxiang Lin, Xinglin Pan, Shaohuai Shi +2
Large language models~(LLMs) are known for their high demand on computing resources and memory due to their substantial model size, which leads to inefficient inference on moderate…