1 paper
Yue Zhang, Yuansheng Chen, Xuan Mo +3
LLM inference serving typically scales out with a two-tier architecture: a cluster router distributes requests to multiple inference engines, each of which then in turn performs it…