large language model evaluation 1preference learning 1query-only supervision 1rubric generation 1synthetic pairwise data 1
From the 1 of 8 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
Haocheng Yang, Licheng Pan, Xiaoxi Li +5
The paper proposes a query‑only method that automatically creates and validates fine‑grained rubrics for evaluating large language models by using synthetic rubric‑conditioned resp…
cs.CL2026
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
Hao Wang, Haocheng Yang, Licheng Pan +7
Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…