1 paper · 1 filter
Seungyeon Jwa, Daechul Ahn, Reokyoung Kim +2
Automatic evaluation with large language models, commonly known as LLM-as-a-judge, is now standard across reasoning and alignment tasks. Despite evaluating many samples in deployme…