91 citations · 355 across the 18 of their papers we have counts for
5 papers · 1 filter
Model-based RL in Contextual Decision Processes: PAC bounds and Exponential Improvements over Model-free Approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy +2
We study the sample complexity of model-based reinforcement learning (henceforth RL) in general contextual decision processes that require strategic exploration to find a near-opti…
Contextual bandits with surrogate losses: Margin bounds and efficient algorithms
Dylan J. Foster, Akshay Krishnamurthy
We use surrogate losses to obtain several new regret bounds and new algorithms for contextual bandit learning. Using the ramp loss, we derive new margin-based regret bounds in term…
Myopic Bayesian Design of Experiments via Posterior Sampling and Probabilistic Programming
Kirthevasan Kandasamy, Willie Neiswanger, Reed Zhang +3
We design a new myopic strategy for a wide class of sequential design of experiment (DOE) problems, where the goal is to collect data in order to to fulfil a certain problem specif…
Semiparametric Contextual Bandits
Akshay Krishnamurthy, Zhiwei Steven Wu, Vasilis Syrgkanis
This paper studies semiparametric contextual bandits, a generalization of the linear stochastic bandit problem where the reward for an action is modeled as a linear function of kno…
On Oracle-Efficient PAC RL with Rich Observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy +3
We study the computational tractability of PAC reinforcement learning with rich observations. We present new provably sample-efficient algorithms for environments with deterministi…