Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Golden Handcuffs make safer AI agents
Aram Ebtekar, Michael K. Cohen
Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward…
cs.LG2024
RL, but don't do anything I wouldn't do
Michael K. Cohen, Marcus Hutter, Yoshua Bengio +1
In reinforcement learning, if the agent's reward differs from the designers' true utility, even only rarely, the state distribution resulting from the agent's policy can be very ba…