works on

From the 1 of 14 linked papers with an AI index.

collaborators

14 papers

cs.CL2026

Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

Haocheng Yang, Licheng Pan, Xiaoxi Li +5

The paper proposes a query‑only method that automatically creates and validates fine‑grained rubrics for evaluating large language models by using synthetic rubric‑conditioned resp…

cs.LG2026

Uncertainty-Aware Reward Modeling for Stable RLHF

Licheng Pan, Haocheng Yang, Haoxuan Li +7

Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…

cs.LG2026

Rethinking the Flow-Based Gradual Domain Adaptation: A Semi-Dual Optimal Transport Perspective

Zhichao Chen, Zhan Zhuang, Yunfei Teng +6

Gradual domain adaptation (GDA) aims to mitigate domain shift by progressively adapting models from the source domain to the target domain via intermediate domains. However, real i…

cs.LG2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Licheng Pan, Haochen Yang, Haoxuan Li +8

Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…

cs.LG2026

DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment

Hao Wang, Licheng Pan, Yuan Lu +7

Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approac…

cs.CL2026

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

Tianqiao Liu, Xueyi Li, Hao Wang +4

Recent advances in large language models (LLMs) have attracted significant interest in extending their capabilities to multimodal scenarios, particularly for speech-to-speech conve…