Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
Zhirui Chen, P. N. Karthik, Yeow Meng Chee +1
We consider a multi-armed bandit setting with finitely many arms, in which each arm yields an -dimensional vector reward upon selection. We assume that the reward of each dimens…
cs.LG2024
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
Zhirui Chen, Vincent Y. F. Tan
We consider the problem of offline reinforcement learning from human feedback (RLHF) with pairwise comparisons proposed by Zhu et al. (2023), where the implicit reward is a linear…