Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Offline Safe Policy Optimization From Heterogeneous Feedback
Ze Gong, Pradeep Varakantham, Akshat Kumar
Offline Preference-based Reinforcement Learning (PbRL) learns rewards and policies aligned with human preferences without the need for extensive reward engineering and direct inter…
cs.AI2025
Safety through feedback in Constrained RL
Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
In safety-critical RL settings, the inclusion of an additional cost function is often favoured over the arduous task of modifying the reward function to ensure the agent's safe beh…