1 paper · 1 filter
Krishnapriya Vishnubhotla, Sowmya Vajjala, Akriti Vij +1
We evaluate the consistency of automated judges in conducting a multi-dimensional safety evaluation in a reference-free setup. Our results indicate that Large Language Models are u…