8 citations · 14 across the 2 of their papers we have counts for
3 papers
cs.LG2021
Improper Reinforcement Learning with Gradient-based Policy Optimization
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan +1
We consider an improper reinforcement learning setting where a learner is given base controllers for an unknown Markov decision process, and wishes to combine them optimally to…
cs.LG2020★ 8 cited
Explicit Best Arm Identification in Linear Bandits Using No-Regret Learners
Mohammadi Zaki, Avi Mohan, Aditya Gopalan
We study the problem of best arm identification in linearly parameterised multi-armed bandits. Given a set of feature vectors a confidence paramet…
cs.LG2019★ 6 cited
Towards Optimal and Efficient Best Arm Identification in Linear Bandits
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan
We give a new algorithm for best arm identification in linearly parameterised bandits in the fixed confidence setting. The algorithm generalises the well-known LUCB algorithm of Ka…