1 paper · 1 filter
Bo Lv, Jingbo Sun
Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Current routing methods primarily…