3 papers
cs.LG2026
-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
Di Wu, Chengshuai Shi, Jing Yang +1
Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone technique for post-training large language models. While most existing approaches rely on the reverse KL-…
cs.LG2025
Greedy Sampling Is Provably Efficient for RLHF
Di Wu, Chengshuai Shi, Jing Yang +1
Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique for post-training large language models. Despite its empirical success, the theoretical understandi…
stat.ML2025
Cost-Aware Optimal Pairwise Pure Exploration
Di Wu, Chengshuai Shi, Ruida Zhou +1
Pure exploration is one of the fundamental problems in multi-armed bandits (MAB). However, existing works mostly focus on specific pure exploration tasks, without a holistic view o…