1 paper
Zifan Lyu, Chahine Nejma, Tobias Wegel +2
Large Language Models are typically benchmarked by evaluating every model on every test query. For practitioners seeking the best model to deploy, this is often wasteful: if a mode…