1 citations · 2 across the 4 of their papers we have counts for
4 papers
Automated Discovery of Functional Actual Causes in Complex Environments
Caleb Chuck, Sankaran Vaidyanathan, Stephen Giguere +3
Reinforcement learning (RL) algorithms often struggle to learn policies that generalize to novel situations due to issues such as causal confusion, overfitting to irrelevant factor…
Learning Optimal Advantage from Preferences and Mistaking it for Reward
W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson +4
We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most re…
Language-guided Task Adaptation for Imitation Learning
Prasoon Goyal, Raymond J. Mooney, Scott Niekum
We introduce a novel setting, wherein an agent needs to learn a task from a demonstration of a related task with the difference between the tasks communicated in natural language.…
SOPE: Spectrum of Off-Policy Estimators
Christina J. Yuan, Yash Chandak, Stephen Giguere +2
Many sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the…