8 citations · 14 across the 3 of their papers we have counts for
4 papers · 1 filter
Better than the Best: Gradient-based Improper Reinforcement Learning for Network Scheduling
Mohammani Zaki, Avi Mohan, Aditya Gopalan +1
We consider the problem of scheduling in constrained queueing networks with a view to minimizing packet delay. Modern communication systems are becoming increasingly complex, and a…
Improper Reinforcement Learning with Gradient-based Policy Optimization
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan +1
We consider an improper reinforcement learning setting where a learner is given base controllers for an unknown Markov decision process, and wishes to combine them optimally to…
Explicit Best Arm Identification in Linear Bandits Using No-Regret Learners
Mohammadi Zaki, Avi Mohan, Aditya Gopalan
We study the problem of best arm identification in linearly parameterised multi-armed bandits. Given a set of feature vectors a confidence paramet…
Towards Optimal and Efficient Best Arm Identification in Linear Bandits
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan
We give a new algorithm for best arm identification in linearly parameterised bandits in the fixed confidence setting. The algorithm generalises the well-known LUCB algorithm of Ka…