Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
Daiki E. Matsunaga, Junho Na, Tri Wahyu Guntara +4
Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centralized Training with Decentralized Execution…
cs.AI2025
Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
Jihwan Jeong, Xiaoyu Wang, Jingmin Wang +2
Offline reinforcement learning (RL) is crucial when online exploration is costly or unsafe but often struggles with high epistemic uncertainty due to limited data. Existing methods…