18 citations · 28 across the 7 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022★ 1 cited
Symbol Guided Hindsight Priors for Reward Learning from Human Preferences
Mudit Verma, Katherine Metcalf
Specifying rewards for reinforcement learned (RL) agents is challenging. Preference-based RL (PbRL) mitigates these challenges by inferring a reward from feedback over sets of traj…
cs.AI2022
Advice Conformance Verification by Reinforcement Learning agents for Human-in-the-Loop
Mudit Verma, Ayush Kharkwal, Subbarao Kambhampati
Human-in-the-loop (HiL) reinforcement learning is gaining traction in domains with large action and state spaces, and sparse rewards by allowing the agent to take advice from HiL.…