3 papers
cs.MA2026
Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces
Dongming Wang, Pengcheng Dai, Wenwu Yu +1
We develop the Continuous Distributed Coupled Policy Gradient (CDCPG) algorithm for cooperative reinforcement learning in networked Markov decision processes with continuous state…
cs.MA2026
Distributed Zeroth-Order Policy Gradient for Networked Multi-agent Reinforcement Learning from Human Feedback
Pengcheng Dai, He Wang, Dongming Wang +2
We study a networked multi-agent reinforcement learning (NMARL) problem with human feedback in an infinite-horizon setting, where agents interact over an underlying network with lo…
cs.MA2025
Distributed scalable coupled policy algorithm for networked multi-agent reinforcement learning
Pengcheng Dai, Dongming Wang, Wenwu Yu +1
This paper studies networked multi-agent reinforcement learning (NMARL) with interdependent rewards and coupled policies. In this setting, each agent's reward depends on its own st…