4 papers
Minimax Optimal Strategy for Delayed Observations in Online Reinforcement Learning
Harin Lee, Kevin Jamieson
We study reinforcement learning with delayed state observation, where the agent observes the current state after some random number of time steps. We propose an algorithm that comb…
Optimal Posterior Sampling for Policy Identification in Tabular Markov Decision Processes
Cyrille Kone, Kevin Jamieson
We study the -PAC policy identification problem in finite-horizon episodic Markov Decision Processes. Existing approaches provide finite-time guarantees for appr…
On the Limitations and Possibilities of Nash Regret Minimization in Zero-Sum Matrix Games under Noisy Feedback
Arnab Maiti, Kevin Jamieson, Lillian J. Ratliff
This paper studies a variant of two-player zero-sum matrix games, where, at each timestep, the row player selects row , the column player selects column , and the row player…
Learning to Actively Learn: A Robust Approach
Jifan Zhang, Lalit Jain, Kevin Jamieson
This work proposes a procedure for designing algorithms for specific adaptive data collection tasks like active learning and pure-exploration multi-armed bandits. Unlike the design…