1 paper · 1 filter
Francisco Eiras, Eliott Zemour, Eric Lin +1
Large Language Model (LLM) based judges form the underpinnings of key safety evaluation processes such as offline benchmarking, automated red-teaming, and online guardrailing. This…