1 paper
Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh
Inference on large-language models (LLMs) is constrained by GPU memory capacity. A sudden increase in the number of inference requests to a cloud-hosted LLM can deplete GPU memory,…