6 papers
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Yaswanth Chittepu, Ativ Joshi, Sohini Chintala +1
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weight updates. However, existing inference-ti…
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee +1
Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost…
Adaptive Margin RLHF via Preference over Preferences
Yaswanth Chittepu, Prasann Singhal, Greg Durrett +1
Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model learning from preferences within Reinfor…
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
Yaswanth Chittepu, Raghavendra Addanki, Tung Mai +2
The development of autonomous machine learning (ML) agents capable of end-to-end data science workflows represents a significant frontier in artificial intelligence. These agents m…
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
Yaswanth Chittepu, Blossom Metevier, Will Schwarzer +3
Existing approaches to language model alignment often treat safety as a tradeoff against helpfulness, which can lead to unacceptable responses in sensitive domains. To ensure relia…
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
Rafael Rafailov, Yaswanth Chittepu, Ryan Park +5
Reinforcement Learning from Human Feedback (RLHF) has been crucial to the recent success of Large Language Models (LLMs), however, it is often a complex and brittle process. In the…