collaborators

5 papers

cs.CL2026

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

Hongli Zhou, Hui Huang, Rui Zhang +5

Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evalu…

cs.CL2026

RM-Distiller: Exploiting Generative LLM for Reward Model Distillation

Hongli Zhou, Hui Huang, Wei Liu +8

Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-quality human preference annotation…

cs.CL2026

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

Hongli Zhou, Hui Huang, Ziqing Zhao +10

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concern…

cs.CL2026

Think-J: Learning to Think for Generative LLM-as-a-Judge

Hui Huang, Yancheng He, Hongli Zhou +5

LLM-as-a-Judge refers to the automatic modeling of preferences for responses generated by Large Language Models (LLMs), which is of significant importance for both LLM evaluation a…

cs.CL2025

An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4

Hui Huang, Xingyuan Bu, Hongli Zhou +5

Recently, there has been a growing trend of utilizing Large Language Model (LLM) to evaluate the quality of other LLMs. Many studies have fine-tuned judge models based on open-sour…