23 citations · 45 across the 9 of their papers we have counts for
20 papers · 1 filter
On Slowly-varying Non-stationary Bandits
Ramakrishnan Krishnamurthy, Aditya Gopalan
We consider minimisation of dynamic regret in non-stationary bandits with a slowly varying property. Namely, we assume that arms' rewards are stochastic and independent over time,…
Better than the Best: Gradient-based Improper Reinforcement Learning for Network Scheduling
Mohammani Zaki, Avi Mohan, Aditya Gopalan +1
We consider the problem of scheduling in constrained queueing networks with a view to minimizing packet delay. Modern communication systems are becoming increasingly complex, and a…
Improper Reinforcement Learning with Gradient-based Policy Optimization
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan +1
We consider an improper reinforcement learning setting where a learner is given base controllers for an unknown Markov decision process, and wishes to combine them optimally to…
Stochastic Linear Bandits with Protected Subspace
Advait Parulekar, Soumya Basu, Aditya Gopalan +2
We study a variant of the stochastic linear bandit problem wherein we optimize a linear objective function but rewards are accrued only orthogonal to an unknown subspace (which we…
No-regret Algorithms for Multi-task Bayesian Optimization
Sayak Ray Chowdhury, Aditya Gopalan
We consider multi-objective optimization (MOO) of an unknown vector-valued function in the non-parametric Bayesian optimization (BO) setting, with the aim being to learn points on…
Explicit Best Arm Identification in Linear Bandits Using No-Regret Learners
Mohammadi Zaki, Avi Mohan, Aditya Gopalan
We study the problem of best arm identification in linearly parameterised multi-armed bandits. Given a set of feature vectors a confidence paramet…