219 citations · 235 across the 2 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2019
SafeLife 1.0: Exploring Side Effects in Complex Environments
Carroll L. Wainwright, Peter Eckersley
We present SafeLife, a publicly available reinforcement learning environment that tests the safety of reinforcement learning agents. It contains complex, dynamic, tunable, procedur…
cs.AI2019★ 16 cited
Impossibility and Uncertainty Theorems in AI Value Alignment (or why your AGI should not have a utility function)
Peter Eckersley
Utility functions or their equivalents (value functions, objective functions, loss functions, reward functions, preference orderings) are a central tool in most current machine lea…