2 papers
cs.AI2026
Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
Andreas Chouliaras, Luke Connolly, Dimitris Chatzpoulos
Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most effective defenses against this are (i) tran…
cs.HC2025
Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
Andreas Chouliaras, Dimitris Chatzopoulos
Reinforcement Learning from Human Feedback (RLHF) relies on preference modeling to align machine learning systems with human values, yet the popular approach of random pair samplin…