1 paper
Tzu-Heng Huang, Shengqi Qiu, Frederic Sala
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalabil…