1 citations · 1 across the 1 of their papers we have counts for
1 paper
Pedro Freire, Adam Gleave, Sam Toyer +1
The objective of many real-world tasks is complex and difficult to procedurally specify. This makes it necessary to use reward or imitation learning algorithms to infer a reward or…