2 papers
cs.LG2025
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
Xiyue Peng, Hengquan Guo, Jiawei Zhang +4
Balancing helpfulness and safety (harmlessness) is a critical challenge in aligning large language models (LLMs). Current approaches often decouple these two objectives, training s…
cs.LG2024
Learning to Schedule Online Tasks with Bandit Feedback
Yongxin Xu, Shangshang Wang, Hengquan Guo +2
Online task scheduling serves an integral role for task-intensive applications in cloud computing and crowdsourcing. Optimal scheduling can enhance system performance, typically me…