From the 1 of 6 linked papers with an AI index.
4 papers · 1 filter
Uncertainty-Aware Reward Modeling for Stable RLHF
Licheng Pan, Haocheng Yang, Haoxuan Li +7
Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing policies to maximize predicted rewards. H…
Looped World Models
Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28
Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive to deploy and prone to compounding error…
Optimal Transport for LLM Reward Modeling from Noisy Preference
Licheng Pan, Haochen Yang, Haoxuan Li +8
Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…
Estimating the Effects of Sample Training Orders for Large Language Models without Retraining
Hao Yang, Haoxuan Li, Mengyue Yang +2
The order of training samples plays a crucial role in large language models (LLMs), significantly impacting both their external performance and internal learning dynamics. Traditio…