10 citations · 10 across the 2 of their papers we have counts for
1 paper · 1 filter
Dario Loi, Elena Maria Muià, Federico Siciliano +4
We present AutoBench, a fully automated and self-sustaining framework for evaluating Large Language Models (LLMs) through reciprocal peer assessment. This paper provides a rigorous…