1 paper
Alessio Ricci Toniolo, Rome Thorstenson, Abinaya Dinesh
Increasingly, LLM inference services proxy client requests to engine replicas distributed globally. Load-balancing policies must jointly account for factors including KV-cache loca…