91 citations · 355 across the 18 of their papers we have counts for
Showing 2016Show all
2 papers · 1 filter
cs.LG2016
Improved Regret Bounds for Oracle-Based Adversarial Contextual Bandits
Vasilis Syrgkanis, Haipeng Luo, Akshay Krishnamurthy +1
We give an oracle-based algorithm for the adversarial contextual bandit problem, where either contexts are drawn i.i.d. or the sequence of contexts is known a priori, but where the…
cs.AI2016
Exploratory Gradient Boosting for Reinforcement Learning in Complex Domains
David Abel, Alekh Agarwal, Fernando Diaz +2
High-dimensional observations and complex real-world dynamics present major challenges in reinforcement learning for both function approximation and exploration. We address both of…