239 citations · 417 across the 34 of their papers we have counts for
5 papers · 1 filter
Instance-optimal PAC Algorithms for Contextual Bandits
Zhaoqi Li, Lillian Ratliff, Houssam Nassif +2
In the stochastic contextual bandit setting, regret-minimizing algorithms have been extensively researched, but their instance-minimizing best-arm identification counterparts remai…
Instance-Dependent Near-Optimal Policy Identification in Linear MDPs via Online Experiment Design
Andrew Wagenmaker, Kevin Jamieson
While much progress has been made in understanding the minimax sample complexity of reinforcement learning (RL) -- the complexity of learning on the "worst-case" instance -- such m…
Active Learning with Safety Constraints
Romain Camilleri, Andrew Wagenmaker, Jamie Morgenstern +2
Active learning methods have shown great promise in reducing the number of samples necessary for learning. As automated learning systems are adopted into real-time, real-world deci…
Active Multi-Task Representation Learning
Yifang Chen, Simon S. Du, Kevin Jamieson
To leverage the power of big data from source tasks and overcome the scarcity of the target task samples, representation learning based on multi-task pretraining has become a stand…
Reward-Free RL is No Harder Than Reward-Aware RL in Linear Markov Decision Processes
Andrew Wagenmaker, Yifang Chen, Max Simchowitz +2
Reward-free reinforcement learning (RL) considers the setting where the agent does not have access to a reward function during exploration, but must propose a near-optimal policy f…