15 citations · 15 across the 2 of their papers we have counts for
2 papers
cs.LG2021
Combining Reward Information from Multiple Sources
Dmitrii Krasheninnikov, Rohin Shah, Herke van Hoof
Given two sources of evidence about a latent variable, one can combine the information from both by multiplying the likelihoods of each piece of evidence. However, when one or both…
cs.LG2019★ 15 cited
Preferences Implicit in the State of the World
Rohin Shah, Dmitrii Krasheninnikov, Jordan Alexander +2
Reinforcement learning (RL) agents optimize only the features specified in a reward function and are indifferent to anything left out inadvertently. This means that we must not onl…