21 citations · 52 across the 3 of their papers we have counts for
3 papers
cs.AI2023★ 21 cited
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Josef Dai, Xuehai Pan, Ruiyang Sun +5
With the development of large language models (LLMs), striking a balance between the performance and safety of AI systems has never been more critical. However, the inherent tensio…
cs.LG2023★ 12 cited
OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research
Jiaming Ji, Jiayi Zhou, Borong Zhang +7
AI systems empowered by reinforcement learning (RL) algorithms harbor the immense potential to catalyze societal advancement, yet their deployment is often impeded by significant s…
cs.LG2022★ 19 cited
Constrained Update Projection Approach to Safe Policy Optimization
Long Yang, Jiaming Ji, Juntao Dai +5
Safe reinforcement learning (RL) studies problems where an intelligent agent has to not only maximize reward but also avoid exploring unsafe areas. In this study, we propose CUP, a…