1 paper · 1 filter
Yulong He, Ivan Smirnov, Dmitry Fedrushkov +2
Effective evaluation of large language models (LLMs) remains a critical bottleneck, as conventional direct scoring often yields inconsistent and opaque judgments. In this work, we…