1 paper · 1 filter
Tao Huang, Pengfei Chen, Kyoka Gong +5
Since the increasing popularity of large language model (LLM) backend systems, it is common and necessary to deploy stable serverless serving of LLM on multi-GPU clusters with auto…