3 papers
cs.LG2026
Unifying Stable Optimization and Reference Regularization in RLHF
Li He, Qiang Qu, He Zhao +4
Reinforcement Learning from Human Feedback (RLHF) has advanced alignment capabilities significantly but remains hindered by two core challenges: \textbf{reward hacking} and \textbf…
cs.AI2025
Direct Advantage Regression: Aligning LLMs with Online AI Reward
Li He, He Zhao, Stephen Wan +3
Online AI Feedback (OAIF) presents a promising alternative to Reinforcement Learning from Human Feedback (RLHF) by utilizing online AI preference in aligning language models (LLMs)…
cs.CL2025
DeFine: A Decomposed and Fine-Grained Annotated Dataset for Long-form Article Generation
Ming Wang, Fang Wang, Minghao Hu +9
Long-form article generation (LFAG) presents challenges such as maintaining logical consistency, comprehensive topic coverage, and narrative coherence across extended articles. Exi…