collaborators

19 papers

cs.AI2026

A Descriptive and Normative Theory of Human Beliefs in RLHF

Sylee Dandekar, Shripad Deshmukh, Frank Chiu +2

Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values. In this work, we propose that human belie…

cs.LG2026

Supervised Reward Inference

Will Schwarzer, Jordan Schneider, Philip S. Thomas +1

Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models. However, humans often indicate their goals throug…

cs.LG2026

Safe Inference-Time Alignment via Lagrangian Reward Augmentation

Yaswanth Chittepu, Ativ Joshi, Sohini Chintala +1

Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-ti…

cs.LG2026

Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control

Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee +1

Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost…

cs.LG2026

Adaptive Margin RLHF via Preference over Preferences

Yaswanth Chittepu, Prasann Singhal, Greg Durrett +1

Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinfor…

cs.AI2026

Hierarchical Experimentalist Agents

Abhranil Chandra, Sankaran Vaidyanathan, Utsav Dhanuka +2

Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-tra…