4 papers
Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
Joanikij Chulev, Hendrik Baier
Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired return command. However, the local mapping…
DiPRL: Learning Discrete Programmatic Policies via Architecture Entropy Regularization
Chengpeng Hu, Yingqian Zhang, Hendrik Baier
Programmatic reinforcement learning (PRL) offers an interpretable alternative to deep reinforcement learning by representing policies as human-readable and -editable programs. Whil…
Scheduling That Speaks: An Interpretable Programmatic Reinforcement Learning Framework
Chengpeng Hu, Yingqian Zhang, Hendrik Baier
Deep reinforcement learning (DRL) has recently emerged as a promising approach to solve combinatorial optimization problems such as job shop scheduling. However, the policies learn…
InnateCoder: Learning Programmatic Options with Foundation Models
Rubens O. Moraes, Quazi Asif Sadmine, Hendrik Baier +1
Outside of transfer learning settings, reinforcement learning agents start their learning process from a clean slate. As a result, such agents have to go through a slow process to…