1 citations · 1 across the 8 of their papers we have counts for
6 papers · 1 filter
Memory Limitations of Prompt Tuning in Transformers
Maxime Meyer, Mario Michelessa, Caroline Chaux +1
Despite the empirical success of prompt tuning in adapting pretrained language models to new tasks, theoretical analyses of its capabilities remain limited. Existing theoretical wo…
Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning
Armin Behnamnia, Gholamali Aminian, Alireza Aghaei +3
Off-policy learning and evaluation leverage logged bandit feedback datasets, which contain context, action, propensity score, and feedback for each data point. These scenarios face…
Best Arm Identification with Possibly Biased Offline Data
Le Yang, Vincent Y. F. Tan, Wang Chi Cheung
We study the best arm identification (BAI) problem with potentially biased offline data in the fixed confidence setting, which commonly arises in real-world scenarios such as clini…
Optimal Multi-Objective Best Arm Identification with Fixed Confidence
Zhirui Chen, P. N. Karthik, Yeow Meng Chee +1
We consider a multi-armed bandit setting with finitely many arms, in which each arm yields an -dimensional vector reward upon selection. We assume that the reward of each dimens…
On the Exponential Convergence for Offline RLHF with Pairwise Comparisons
Zhirui Chen, Vincent Y. F. Tan
We consider the problem of offline reinforcement learning from human feedback (RLHF) with pairwise comparisons proposed by Zhu et al. (2023), where the implicit reward is a linear…
Fixed-Budget Differentially Private Best Arm Identification
Zhirui Chen, P. N. Karthik, Yeow Meng Chee +1
We study best arm identification (BAI) in linear bandits in the fixed-budget regime under differential privacy constraints, when the arm rewards are supported on the unit interval.…