6 papers
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
Zhirui Chen, Vincent Y. F. Tan
We consider the problem of offline reinforcement learning from human feedback (RLHF) with pairwise comparisons proposed by Zhu et al. (2023), where the implicit reward is a linear…
Finite-Time Minimax Bounds and an Optimal Lyapunov Policy in Queueing Control
Yujie Liu, Vincent Y. F. Tan, Yunbei Xu
We introduce an original minimax framework for finite-time performance analysis in queueing control and propose a surprisingly simple Lyapunov-based scheduling policy with superior…
Memory Limitations of Prompt Tuning in Transformers
Maxime Meyer, Mario Michelessa, Caroline Chaux +1
Despite the empirical success of prompt tuning in adapting pretrained language models to new tasks, theoretical analyses of its capabilities remain limited. Existing theoretical wo…
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
Armin Behnamnia, Gholamali Aminian, Alireza Aghaei +3
Off-policy learning and evaluation leverage logged bandit feedback datasets, which contain context, action, propensity score, and feedback for each data point. These scenarios face…
Best Arm Identification with Possibly Biased Offline Data
Le Yang, Vincent Y. F. Tan, Wang Chi Cheung
We study the best arm identification (BAI) problem with potentially biased offline data in the fixed confidence setting, which commonly arises in real-world scenarios such as clini…
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
Zhirui Chen, P. N. Karthik, Yeow Meng Chee +1
We consider a multi-armed bandit setting with finitely many arms, in which each arm yields an -dimensional vector reward upon selection. We assume that the reward of each dimens…