1 paper
Anissa Alloula, Federico Licini, Ava Batchkala +1
LLMs-as-judges are the only way to evaluate safety at scale. Despite their importance, LLM-judges themselves are rarely evaluated beyond human agreement in simple, static benchmark…