1 paper
Leon Eshuijs, Archie Chaudhury, Alan McBeth +1
LLM-as-a-judge is widely used as a scalable substitute for human evaluation, yet current approaches rely on black-box access and struggle to detect subtle dishonesty, such as sycop…