3 papers
cs.LG2025
Learning to Answer from Correct Demonstrations
Nirmit Joshi, Gene Li, Siddharth Bhandari +3
We study the problem of learning to generate an answer (or completion) to a question (or prompt), where there could be multiple correct answers, any one of which is acceptable at t…
cs.LG2025
Agnostic Reinforcement Learning: Foundations and Algorithms
Gene Li
Reinforcement Learning (RL) has demonstrated tremendous empirical success across numerous challenging domains. However, we lack a strong theoretical understanding of the statistica…
cs.LG2025
The Role of Environment Access in Agnostic Reinforcement Learning
Akshay Krishnamurthy, Gene Li, Ayush Sekhari
We study Reinforcement Learning (RL) in environments with large state spaces, where function approximation is required for sample-efficient learning. Departing from a long history…