1 paper
Dachuan Lin, Guobin Shen, Zihao Yang +3
Safety evaluation of large language models (LLMs) increasingly relies on LLM-as-a-judge pipelines, but strong judges can still be expensive to use at scale. We study whether struct…