1 paper
Delip Rao, Chris Callison-Burch
Whether a rubric-based LLM judge can replace human annotation is decided by its measured agreement with human labels. Yet the same verdicts can support wildly varying agreement num…