collaborators

8 papers

cs.CL2026

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

Haocheng Yang, Licheng Pan, Xiaoxi Li +5

The paper proposes a query‑only method that automatically creates and validates fine‑grained rubrics for evaluating large language models by using synthetic rubric‑conditioned resp…

cs.LG2026

Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

Ran Piao, Tsai-Ning Wang, Martijn den Dekker +4

Clinical audio diagnosis in low-resource settings requires models that identify conditions from minimal examples without large annotated corpora. We propose Federated Self-Contextu…

cs.LG2026

Uncertainty-Aware Reward Modeling for Stable RLHF

Licheng Pan, Haocheng Yang, Haoxuan Li +7

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…

cs.CL2026

Evaluating Chinese Ambiguity Understanding in Large Language Models

Junwen Mo, Yuanzhi Lu, Yifang Xue +2

Linguistic ambiguity is critical to the robustness of Large Language Models (LLMs), yet existing research focuses mostly on English, with limited attention devoted to Chinese. Exis…

cs.LG2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Licheng Pan, Haochen Yang, Haoxuan Li +8

Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…

cs.CL2026

ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment

Hao Wang, Haocheng Yang, Licheng Pan +7

Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…