6 papers
Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction
Xingguo Chen, Zhiang He, Yuchen Shen +4
Temporal-difference learning with function approximation can be unstable under off-policy sampling. TDC stabilizes off-policy TD through an auxiliary covariance correction, and TDR…
Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction
Xingguo Chen, Yuchen Shen, Shangdong Yang +3
Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongly affected by the geometry i…
Fully Decentralized Cooperative Multi-Agent Reinforcement Learning is A Context Modeling Problem
Chao Li, Bingkun Bao, Yang Gao
This paper studies fully decentralized cooperative multi-agent reinforcement learning, where each agent solely observes the states, its local actions, and the shared rewards. The i…
Regularized Centered Emphatic Temporal Difference Learning
Xingguo Chen, Chaohui Wu, Jinguo Ye +5
Off-policy temporal-difference (TD) learning with function approximation faces a structural tradeoff among stability, projection geometry, and variance control. Emphatic TD (ETD) i…
Bitboard version of Tetris AI
Xingguo Chen, Pingshou Xiong, Zhenyu Luo +6
The efficiency of game engines and policy optimization algorithms is crucial for training reinforcement learning (RL) agents in complex sequential decision-making tasks, such as Te…
OpenGuanDan: A Large-Scale Imperfect Information Game Benchmark
Chao Li, Shangdong Yang, Chiheng Zhan +5
The advancement of data-driven artificial intelligence (AI), particularly machine learning, heavily depends on large-scale benchmarks. Despite remarkable progress across domains ra…