12 citations · 25 across the 5 of their papers we have counts for
4 papers · 1 filter
Dynamic probabilistic logic models for effective abstractions in RL
Harsha Kokel, Arjun Manoharan, Sriraam Natarajan +2
State abstraction enables sample-efficient learning and better task transfer in complex reinforcement learning environments. Recently, we proposed RePReL (Kokel et al. 2021), a hie…
Avoiding Side Effects in Complex Environments
Alexander Matt Turner, Neale Ratzlaff, Prasad Tadepalli
Reward function specification can be difficult. Rewarding the agent for making a widget may be easy, but penalizing the multitude of possible negative side effects is hard. In toy…
The Choice Function Framework for Online Policy Improvement
Murugeswari Issakkimuthu, Alan Fern, Prasad Tadepalli
There are notable examples of online search improving over hand-coded or learned policies (e.g. AlphaZero) for sequential decision making. It is not clear, however, whether or not…
Conservative Agency via Attainable Utility Preservation
Alexander Matt Turner, Dylan Hadfield-Menell, Prasad Tadepalli
Reward functions are easy to misspecify; although designers can make corrections after observing mistakes, an agent pursuing a misspecified reward function can irreversibly change…