1 paper · 1 filter
Manan Gupta, Inderjeet Nair, Lu Wang +1
The LLM-as-a-judge paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: that judges evaluate text st…