1 paper
Haolong Chen, Zhengyuan Xin, Liang Zhang +2
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model pe…