3 citations · 3 across the 1 of their papers we have counts for
1 paper
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson +2
Learning policies via preference-based reward learning is an increasingly popular method for customizing agent behavior, but has been shown anecdotally to be prone to spurious corr…