activity
20172022
most citedOptimal Regret Algorithm for Pseudo-1d Bandit Convex Optimization

2 citations · 2 across the 12 of their papers we have counts for

collaborators

20 papers

cs.LG2022

One Arrow, Two Kills: An Unified Framework for Achieving Optimal Regret Guarantees in Sleeping Bandits

Pierre Gaillard, Aadirupa Saha, Soham Dan

We address the problem of \emph{`Internal Regret'} in \emph{Sleeping Bandits} in the fully adversarial setup, as well as draw connections between different existing notions of slee…

cs.LG2022

ANACONDA: An Improved Dynamic Regret Algorithm for Adaptive Non-Stationary Dueling Bandits

Thomas Kleine Buening, Aadirupa Saha

We study the problem of non-stationary dueling bandits and provide the first adaptive dynamic regret algorithm for this problem. The only two existing attempts in this line of work…

math.OC2022

Dueling Convex Optimization with General Preferences

Aadirupa Saha, Tomer Koren, Yishay Mansour

We address the problem of \emph{convex optimization with dueling feedback}, where the goal is to minimize a convex function given a weaker form of \emph{dueling} feedback. Each que…

cs.LG2022

Exploiting Correlation to Achieve Faster Learning Rates in Low-Rank Preference Bandits

Suprovat Ghoshal, Aadirupa Saha

We introduce the \emph{Correlated Preference Bandits} problem with random utility-based choice models (RUMs), where the goal is to identify the best item from a given pool of i…

cs.LG2022

Versatile Dueling Bandits: Best-of-both-World Analyses for Online Learning from Preferences

Aadirupa Saha, Pierre Gaillard

We study the problem of -armed dueling bandit for both stochastic and adversarial environments, where the goal of the learner is to aggregate information through relative prefer…

cs.LG2021

Strategically Efficient Exploration in Competitive Multi-agent Reinforcement Learning

Robert Loftin, Aadirupa Saha, Sam Devlin +1

High sample complexity remains a barrier to the application of reinforcement learning (RL), particularly in multi-agent systems. A large body of work has demonstrated that explorat…