1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 1 cited
Learning Optimal Advantage from Preferences and Mistaking it for Reward
W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson +4
We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most re…
cs.AI2022
BRTDP: A Belief Branch and Bound Real-Time Dynamic Programming Approach to Solving POMDPs
Sigurdur Orn Adalgeirsson, Cynthia Breazeal
Partially Observable Markov Decision Processes (POMDPs) offer a promising world representation for autonomous agents, as they can model both transitional and perceptual uncertainti…