5 papers
HiMAP-Travel: Hierarchical Multi-Agent Planning for Long-Horizon Constrained Travel
The Viet Bui, Wenjun Li, Yong Liu
Sequential LLM agents fail on long-horizon planning with hard constraints like budgets and diversity requirements. As planning progresses and context grows, these agents drift from…
Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning
The Viet Bui, Tien Mai, Hong Thanh Nguyen
We study the problem of online multi-agent reinforcement learning (MARL) in environments with sparse rewards, where reward feedback is not provided at each interaction but only rev…
MisoDICE: Multi-Agent Imitation from Unlabeled Mixed-Quality Demonstrations
The Viet Bui, Tien Mai, Hong Thanh Nguyen
We study offline imitation learning (IL) in cooperative multi-agent settings, where demonstrations have unlabeled mixed quality - containing both expert and suboptimal trajectories…
O-MAPL: Offline Multi-agent Preference Learning
The Viet Bui, Tien Mai, Hong Thanh Nguyen
Inferring reward functions from demonstrations is a key challenge in reinforcement learning (RL), particularly in multi-agent RL (MARL), where large joint state-action spaces and c…
ComaDICE: Offline Cooperative Multi-Agent Reinforcement Learning with Stationary Distribution Shift Regularization
The Viet Bui, Thanh Hong Nguyen, Tien Mai
Offline reinforcement learning (RL) has garnered significant attention for its ability to learn effective policies from pre-collected datasets without the need for further environm…