2 papers
cs.LG2026
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Jonathan Colaço Carr, Jonathan Colaço Carr, Prakash Panangaden +2
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to…
cs.LG2025
Focused Skill Discovery: Learning to Control Specific State Variables while Minimizing Side Effects
Jonathan Colaço Carr, Qinyi Sun, Cameron Allen
Skills are essential for unlocking higher levels of problem solving. A common approach to discovering these skills is to learn ones that reliably reach different states, thus empow…