3 papers
cs.CL2025
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
Hongli Zhou, Hui Huang, Ziqing Zhao +10
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concern…
cs.CL2025
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training
Hui Huang, Jiaheng Liu, Yancheng He +5
Complex instruction-following with elaborate constraints is imperative for Large Language Models (LLMs). While existing methods have constructed data for complex instruction alignm…
cs.CL2024
Mitigating the Bias of Large Language Model Evaluation
Hongli Zhou, Hui Huang, Yunfei Long +5
Recently, there has been a trend of evaluating the Large Language Model (LLM) quality in the flavor of LLM-as-a-Judge, namely leveraging another LLM to evaluate the current output…