Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Can Differentiable Decision Trees Enable Interpretable Reward Learning from Human Feedback?
Akansha Kalra, Daniel S. Brown
Reinforcement Learning from Human Feedback (RLHF) has emerged as a popular paradigm for capturing human intent to alleviate the challenges of hand-crafting the reward values. Despi…
cs.LG2024
Bayesian Robust Optimization for Imitation Learning
Daniel S. Brown, Scott Niekum, Marek Petrik
One of the main challenges in imitation learning is determining what action an agent should take when outside the state distribution of the demonstrations. Inverse reinforcement le…