1 paper
Pratik S. Sachdeva, Nathan Boudol
LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as…