2 papers
cs.LG2025
Fusing Rewards and Preferences in Reinforcement Learning
Sadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash +1
We present Dual-Feedback Actor (DFA), a reinforcement learning algorithm that fuses both individual rewards and pairwise preferences (if available) into a single update rule. DFA u…
cs.LG2025
Hierarchical Reinforcement Learning with Targeted Causal Interventions
Sadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash +1
Hierarchical reinforcement learning (HRL) improves the efficiency of long-horizon reinforcement-learning tasks with sparse rewards by decomposing the task into a hierarchy of subgo…