collaborators

6 papers

cs.AI2026

Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction

Xingguo Chen, Zhiang He, Yuchen Shen +4

Temporal-difference learning with function approximation can be unstable under off-policy sampling. TDC stabilizes off-policy TD through an auxiliary covariance correction, and TDR…

cs.AI2026

Behavior-Induced Mirror-Prox Temporal-Difference Learning for Faster Off-Policy Prediction

Xingguo Chen, Yuchen Shen, Shangdong Yang +3

Gradient temporal-difference methods provide stable off-policy prediction with linear function approximation, but their practical performance is strongly affected by the geometry i…

cs.LG2026

Fully Decentralized Cooperative Multi-Agent Reinforcement Learning is A Context Modeling Problem

Chao Li, Bingkun Bao, Yang Gao

This paper studies fully decentralized cooperative multi-agent reinforcement learning, where each agent solely observes the states, its local actions, and the shared rewards. The i…

cs.AI2026

Regularized Centered Emphatic Temporal Difference Learning

Xingguo Chen, Chaohui Wu, Jinguo Ye +5

Off-policy temporal-difference (TD) learning with function approximation faces a structural tradeoff among stability, projection geometry, and variance control. Emphatic TD (ETD) i…

cs.AI2026

Bitboard version of Tetris AI

Xingguo Chen, Pingshou Xiong, Zhenyu Luo +6

The efficiency of game engines and policy optimization algorithms is crucial for training reinforcement learning (RL) agents in complex sequential decision-making tasks, such as Te…

cs.AI2026

OpenGuanDan: A Large-Scale Imperfect Information Game Benchmark

Chao Li, Shangdong Yang, Chiheng Zhan +5

The advancement of data-driven artificial intelligence (AI), particularly machine learning, heavily depends on large-scale benchmarks. Despite remarkable progress across domains ra…