6 papers
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Post-training LLMs with RLHF and preference optimization methods (e.g., DPO, IPO) has greatly improved alignment, yet these approaches assume a single objective. In reality, humans…
Best Policy Learning from Trajectory Preference Feedback
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Reinforcement Learning from Human Feedback (RLHF) has emerged as a powerful approach for aligning generative models, but its reliance on learned reward models makes it vulnerable t…
Multi-Objective Reward and Preference Optimization: Theory and Algorithms
Akhil Agnihotri
This thesis develops theoretical frameworks and algorithms that advance constrained reinforcement learning (RL) across control, preference learning, and alignment of large language…
Online Bandit Learning with Offline Preference Data for Improved RLHF
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or p…
e-COP : Episodic Constrained Optimization of Policies
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
In this paper, we present the algorithm, the first policy optimization algorithm for constrained Reinforcement Learning (RL) in episodic (finite horizon) settings.…
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
Akhil Agnihotri, Rahul Jain, Haipeng Luo
Reinforcement Learning (RL) for constrained MDPs (CMDPs) is an increasingly important problem for various applications. Often, the average criterion is more suitable than the disco…