1 paper
Dongjie Xu, Julius, Hanchi Dong +6
Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three structural limitations: data distr…