1 paper
Yangsibo Huang, Milad Nasr, Anastasios Angelopoulos +10
It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or ski…