2 citations · 3 across the 4 of their papers we have counts for
4 papers
Methods and Mechanisms for Interactive Novelty Handling in Adversarial Environments
Tung Thai, Ming Shen, Mayank Garg +10
Learning to detect, characterize and accommodate novelties is a challenge that agents operating in open-world domains need to address to be able to guarantee satisfactory task perf…
Exploiting Unlabeled Data for Feedback Efficient Human Preference based Reinforcement Learning
Mudit Verma, Siddhant Bhambri, Subbarao Kambhampati
Preference Based Reinforcement Learning has shown much promise for utilizing human binary feedback on queried trajectory pairs to recover the underlying reward model of the Human i…
A State Augmentation based approach to Reinforcement Learning from Human Preferences
Mudit Verma, Subbarao Kambhampati
Reinforcement Learning has suffered from poor reward specification, and issues for reward hacking even in simple enough domains. Preference Based Reinforcement Learning attempts to…
Data Driven Reward Initialization for Preference based Reinforcement Learning
Mudit Verma, Subbarao Kambhampati
Preference-based Reinforcement Learning (PbRL) methods utilize binary feedback from the human in the loop (HiL) over queried trajectory pairs to learn a reward model in an attempt…