19 papers
A Descriptive and Normative Theory of Human Beliefs in RLHF
Sylee Dandekar, Shripad Deshmukh, Frank Chiu +2
Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human belie…
Supervised Reward Inference
Will Schwarzer, Jordan Schneider, Philip S. Thomas +1
Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models. However, humans often indicate their goals throug…
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Yaswanth Chittepu, Ativ Joshi, Sohini Chintala +1
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-ti…
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee +1
Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost…
Adaptive Margin RLHF via Preference over Preferences
Yaswanth Chittepu, Prasann Singhal, Greg Durrett +1
Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinfor…
Hierarchical Experimentalist Agents
Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka +2
Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-tra…