1 paper · 1 filter
Junhyuk Choi, Sohhyung Park, Chanhee Cho +2
While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether…