2 papers
cs.LG2026
AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin +6
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected base…
cs.LG2026
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
Zeyu Cao, Xuan Guo, Cheng Zhang +3
As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates whether these retired GPUs can find a produ…