2 papers
cs.LG2023
Kernelized Offline Contextual Dueling Bandits
Viraj Mehta, Ojash Neopane, Vikramjeet Das +3
Preference-based feedback is important for many applications where direct evaluation of a reward function is not feasible. A notable recent example arises in reinforcement learning…
cs.LG2021
Best Arm Identification under Additive Transfer Bandits
Ojash Neopane, Aaditya Ramdas, Aarti Singh
We consider a variant of the best arm identification (BAI) problem in multi-armed bandits (MAB) in which there are two sets of arms (source and target), and the objective is to det…