2 papers
cs.LG2026
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
Jonaid Shianifar, Michael Schukat, Karl Mason
Multi-objective reinforcement learning (MORL) enables agents to optimize vector-valued rewards while respecting user preferences. CAPQL, a preference-conditioned actor-critic metho…
cs.LG2025
Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
Xue Yang, Michael Schukat, Junlin Lu +3
Reinforcement learning (RL) excels in various applications but struggles in dynamic environments where the underlying Markov decision process evolves. Continual reinforcement learn…