1 paper
Borja Ibarz, Jan Leike, Tobias Pohlen +3
To solve complex real-world problems with reinforcement learning, we cannot rely on manually specified reward functions. Instead, we can have humans communicate an objective to the…