Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
Katherine M. Collins, Najoung Kim, Yonatan Bitton +15
Human feedback plays a critical role in learning and refining reward models for text-to-image generation, but the optimal form the feedback should take for learning an accurate rew…
cs.LG2024
e-COP : Episodic Constrained Optimization of Policies
Akhil Agnihotri, Rahul Jain, Deepak Ramachandran +1
In this paper, we present the algorithm, the first policy optimization algorithm for constrained Reinforcement Learning (RL) in episodic (finite horizon) settings.…