1 paper
Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin +6
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected base…