5 papers
The Sample Complexity of Policy Learning with Mu-Resets
Gene Li
We study policy-based reinforcement learning under the -resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajec…
Learning to Answer from Correct Demonstrations
Nirmit Joshi, Gene Li, Siddharth Bhandari +3
We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at t…
Agnostic Reinforcement Learning: Foundations and Algorithms
Gene Li
Reinforcement Learning (RL) has demonstrated tremendous empirical success across numerous challenging domains. However, we lack a strong theoretical understanding of the statistica…
The Role of Environment Access in Agnostic Reinforcement Learning
Akshay Krishnamurthy, Gene Li, Ayush Sekhari
We study Reinforcement Learning (RL) in environments with large state spaces, where function approximation is required for sample-efficient learning. Departing from a long history…
Reinforcement Learning with Intrinsically Motivated Feedback Graph for Lost-sales Inventory Control
Zifan Liu, Xinran Li, Shibo Chen +3
Reinforcement learning (RL) has proven to be well-performed and general-purpose in the inventory control (IC). However, further improvement of RL algorithms in the IC domain is imp…