1 paper · 1 filter
Mingyuan Xu, Xinzi Tan, Jiawei Wu +1
Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is…