1 paper
Amey Agrawal, Mayank Yadav, Sukrit Kumar +9
Deploying LLMs efficiently requires testing hundreds of serving configurations, but evaluating each one on a GPU cluster takes hours and costs thousands of dollars. Discrete-event…