1 paper · 1 filter
Anissa Alloula, Federico Licini, Ava Batchkala +1
LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmark…