1 citations · 1 across the 7 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.LG2024
Online Bandit Learning with Offline Preference Data for Improved RLHF
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
Reinforcement Learning with Human Feedback (RLHF) is at the core of fine-tuning methods for generative AI models for language and images. Such feedback is often sought as rank or p…
cs.LG2024
e-COP : Episodic Constrained Optimization of Policies
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
In this paper, we present the algorithm, the first policy optimization algorithm for constrained Reinforcement Learning (RL) in episodic (finite horizon) settings.…