7 papers
Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning
The Viet Bui, Tien Mai, Hong Thanh Nguyen
We study the problem of online multi-agent reinforcement learning (MARL) in environments with sparse rewards, where reward feedback is not provided at each interaction but only rev…
Learning What to Do and What Not To Do: Offline Imitation from Expert and Undesirable Demonstrations
Huy Hoang, Tien Mai, Pradeep Varakantham +1
Offline imitation learning typically learns from expert and unlabeled demonstrations, yet often overlooks the valuable signal in explicitly undesirable behaviors. In this work, we…
MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations
The Viet Bui, Tien Mai, Hong Thanh Nguyen
We study offline imitation learning (IL) in cooperative multi-agent settings, where demonstrations have unlabeled mixed quality - containing both expert and suboptimal trajectories…
O-MAPL: Offline Multi-agent Preference Learning
The Viet Bui, Tien Mai, Hong Thanh Nguyen
Inferring reward functions from demonstrations is a key challenge in reinforcement learning (RL), particularly in multi-agent RL (MARL), where large joint state-action spaces and c…
UNIQ: Offline Inverse Q-learning for Avoiding Undesirable Demonstrations
Huy Hoang, Tien Mai, Pradeep Varakantham
We address the problem of offline learning a policy that avoids undesirable demonstrations. Unlike conventional offline imitation learning approaches that aim to imitate expert or…
SPRINQL: Sub-optimal Demonstrations driven Offline Imitation Learning
Huy Hoang, Tien Mai, Pradeep Varakantham
We focus on offline imitation learning (IL), which aims to mimic an expert's behavior using demonstrations without any interaction with the environment. One of the main challenges…