1 paper
Andi Peng, Yuying Sun, Tianmin Shu +1
Humans use social context to specify preferences over behaviors, i.e. their reward functions. Yet, algorithms for inferring reward models from preference data do not take this soci…