1 paper
Tanmay Ambadkar, Sourav Panda, Shreyash Kale +2
Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalabl…