1 paper
Azam Ikram, Xiang Li, Sameh Elnikety +1
The rapid advancement of Large Language Models (LLMs) has driven the need for more efficient serving strategies. In this context, efficiency refers to the proportion of requests th…