3 papers
cs.LG2026
Quantifying Potential Observation Missingness in Inverse Reinforcement Learning
Leo Benac, Abhishek Sharma, Alihan Huyuk +1
Inverse reinforcement learning (IRL), which infers reward functions from demonstrations, is a valuable tool for modeling and understanding decision-making behavior. Many variants o…
cs.LG2026
Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
Leo Benac, Abhishek Sharma, Sonali Parbhoo +1
We consider the problem of estimating the transition dynamics from near-optimal expert trajectories in the context of offline model-based reinforcement learning. We develop a…
cs.LG2024
Decision-Point Guided Safe Policy Improvement
Abhishek Sharma, Leo Benac, Sonali Parbhoo +1
Within batch reinforcement learning, safe policy improvement (SPI) seeks to ensure that the learnt policy performs at least as well as the behavior policy that generated the datase…