4 citations · 13 across the 7 of their papers we have counts for
11 papers
Collaborative Multi-Agent Heterogeneous Multi-Armed Bandits
Ronshee Chawla, Daniel Vial, Sanjay Shakkottai +1
The study of collaborative multi-agent bandits has attracted significant attention recently. In light of this, we initiate the study of a new collaborative setting, consisting of $…
Minimax Regret for Cascading Bandits
Daniel Vial, Sujay Sanghavi, Sanjay Shakkottai +1
Cascading bandits is a natural and popular model that frames the task of learning to rank from Bernoulli click feedback in a bandit setting. For the case of unstructured rewards, w…
Robust Multi-Agent Bandits Over Undirected Graphs
Daniel Vial, Sanjay Shakkottai, R. Srikant
We consider a multi-agent multi-armed bandit setting in which honest agents collaborate over a network to minimize regret but malicious agents can disrupt learning arbitrar…
Improved Algorithms for Misspecified Linear Markov Decision Processes
Daniel Vial, Advait Parulekar, Sanjay Shakkottai +1
For the misspecified linear Markov decision process (MLMDP) model of Jin et al. [2020], we propose an algorithm with three desirable properties. (P1) Its regret after episodes…
Regret Bounds for Stochastic Shortest Path Problems with Linear Function Approximation
Daniel Vial, Advait Parulekar, Sanjay Shakkottai +1
We propose an algorithm that uses linear function approximation (LFA) for stochastic shortest path (SSP). Under minimal assumptions, it obtains sublinear regret, is computationally…
One-bit feedback is sufficient for upper confidence bound policies
Daniel Vial, Sanjay Shakkottai, R. Srikant
We consider a variant of the traditional multi-armed bandit problem in which each arm is only able to provide one-bit feedback during each pull based on its past history of rewards…