1 paper
Yao Fu, Leyang Xue, Yeqi Huang +4
This paper presents ServerlessLLM, a distributed system designed to support low-latency serverless inference for Large Language Models (LLMs). By harnessing the substantial near-GP…