1 paper
Yukyung Lee, Joonghoon Kim, Jaehee Kim +4
Existing LLM-as-a-Judge approaches for evaluating text generation suffer from rating inconsistencies, with low agreement and high rating variance across different evaluator models.…