5 papers
Offline Two-Player Zero-Sum Markov Games with KL Regularization
Claire Chen, Yuheng Zhang, Xinyu Liu +3
We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shi…
Beyond Pessimism: Offline Learning in KL-regularized Games
Yuheng Zhang, Claire Chen, Nan Jiang
We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized with respect to a fixed reference policy through KL regularization. Prior work r…
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
Xiang Li, Yuheng Zhang, Nan Jiang
We investigate the theoretical aspects of offline reinforcement learning (RL) under general function approximation. While prior works (e.g., Xie et al., 2021) have established the…
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
Yuheng Zhang, Mingyue Huo, Minghao Zhu +2
Reward models (RMs) are widely used as optimization targets in reinforcement learning from human feedback (RLHF), yet they remain vulnerable to reward hacking. Existing attacks mai…
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
Yuheng Zhang, Nan Jiang
We investigate off-policy evaluation (OPE), a central and fundamental problem in reinforcement learning (RL), in the challenging setting of Partially Observable Markov Decision Pro…