1 paper · 1 filter
Daniel Lee, Harsh Sharma, Eunkyu Park +5
MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heterogeneous. We introduce VOIR DIR…