4.3k citations · 4.3k across the 2 of their papers we have counts for
3 papers
cs.CL2022★ 4.3k cited
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang +17
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic,…
cs.LG2020
Safety Aware Reinforcement Learning (SARL)
Santiago Miret, Somdeb Majumdar, Carroll Wainwright
As reinforcement learning agents become increasingly integrated into complex, real-world environments, designing for safety becomes a critical consideration. We specifically focus…
cs.AI2019
SafeLife 1.0: Exploring Side Effects in Complex Environments
Carroll L. Wainwright, Peter Eckersley
We present SafeLife, a publicly available reinforcement learning environment that tests the safety of reinforcement learning agents. It contains complex, dynamic, tunable, procedur…