Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
Zhirui Chen, Vincent Y. F. Tan
We consider the problem of offline reinforcement learning from human feedback (RLHF) with pairwise comparisons proposed by Zhu et al. (2023), where the implicit reward is a linear…
cs.LG2025
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
Zhirui Chen, P. N. Karthik, Yeow Meng Chee +1
We consider a multi-armed bandit setting with finitely many arms, in which each arm yields an -dimensional vector reward upon selection. We assume that the reward of each dimens…