1 paper · 1 filter
Octavian M. Machidon, Alina L. Machidon, Vojko Strahovnik +3
Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and…