2 papers
cs.LG2024
DP-Dueling: Learning from Preference Feedback without Compromising User Privacy
Aadirupa Saha, Hilal Asi
We consider the well-studied dueling bandit problem, where a learner aims to identify near-optimal actions using pairwise comparisons, under the constraint of differential privacy.…
cs.DS2023
Dueling Optimization with a Monotone Adversary
Avrim Blum, Meghal Gupta, Gene Li +3
We introduce and study the problem of dueling optimization with a monotone adversary, which is a generalization of (noiseless) dueling convex optimization. The goal is to design an…