3 papers
cs.LG2026
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Jonathan Colaço Carr, Prakash Panangaden, Doina Precup +1
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…
cs.LG2025
Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects
Jonathan Colaço Carr, Qinyi Sun, Cameron Allen
Skills are essential for unlocking higher levels of problem solving. A common approach to discovering these skills is to learn ones that reliably reach different states, thus empow…
cs.LG2023
Conditions on Preference Relations that Guarantee the Existence of Optimal Policies
Jonathan Colaço Carr, Prakash Panangaden, Doina Precup
Learning from Preferential Feedback (LfPF) plays an essential role in training Large Language Models, as well as certain types of interactive learning agents. However, a substantia…