3 papers
cs.CL2025
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
Hongli Zhou, Hui Huang, Ziqing Zhao +10
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concern…
cs.CL2025
Think-J: Learning to Think for Generative LLM-as-a-Judge
Hui Huang, Yancheng He, Hongli Zhou +5
LLM-as-a-Judge refers to the automatic modeling of preferences for responses generated by Large Language Models (LLMs), which is of significant importance for both LLM evaluation a…
cs.CL2024
Mitigating the Bias of Large Language Model Evaluation
Hongli Zhou, Hui Huang, Yunfei Long +5
Recently, there has been a trend of evaluating the Large Language Model (LLM) quality in the flavor of LLM-as-a-Judge, namely leveraging another LLM to evaluate the current output…