5 citations · 5 across the 3 of their papers we have counts for
1 paper · 1 filter
Yujun Zhou, Yue Huang, Han Bao +8
While most AI alignment research focuses on preventing models from generating explicitly harmful content, a more subtle risk arises from capability-seeking RL training in vulnerabl…