Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Yaswanth Chittepu, Blossom Metevier, Will Schwarzer +3
Existing approaches to language model alignment often treat safety as a tradeoff against helpfulness, which can lead to unacceptable responses in sensitive domains. To ensure relia…
cs.LG2025
Supervised Reward Inference
Will Schwarzer, Jordan Schneider, Philip S. Thomas +1
Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models. However, humans often indicate their goals throug…