collaborators

6 papers

cs.CL2026

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

Haocheng Yang, Licheng Pan, Xiaoxi Li +5

The paper proposes a query‑only method that automatically creates and validates fine‑grained rubrics for evaluating large language models by using synthetic rubric‑conditioned resp…

cs.LG2026

Uncertainty-Aware Reward Modeling for Stable RLHF

Licheng Pan, Haocheng Yang, Haoxuan Li +7

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…

cs.LG2026

Looped World Models

Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28

Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive to deploy and prone to compounding error…

cs.LG2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Licheng Pan, Haochen Yang, Haoxuan Li +8

Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…

cs.CL2025

Mitigating Hidden Confounding by Progressive Confounder Imputation via Large Language Models

Hao Yang, Haoxuan Li, Luyu Chen +3

Hidden confounding remains a central challenge in estimating treatment effects from observational data, as unobserved variables can lead to biased causal estimates. While recent wo…

cs.LG2025

Estimating the Effects of Sample Training Orders for Large Language Models without Retraining

Hao Yang, Haoxuan Li, Mengyue Yang +2

The order of training samples plays a crucial role in large language models (LLMs), significantly impacting both their external performance and internal learning dynamics. Traditio…