6 papers
Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics
Zuyuan Zhang, Sizhe Tang, Tian Lan
Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the cen…
Structuring Value Representations via Geometric Coherence in Markov Decision Processes
Zuyuan Zhang, Zeyu Fang, Tian Lan
Geometric properties can be leveraged to stabilize and speed reinforcement learning. Existing examples include encoding symmetry structure, geometry-aware data augmentation, and en…
Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning
Zeyu Fang, Zuyuan Zhang, Mahdi Imani +1
Model-based offline reinforcement learning is brittle under distribution shift: policy improvement drives rollouts into state--action regions weakly supported by the dataset, where…
Geometry of Drifting MDPs with Path-Integral Stability Certificates
Zuyuan Zhang, Mahdi Imani, Tian Lan
Real-world reinforcement learning is often \emph{nonstationary}: rewards and dynamics drift, accelerate, oscillate, and trigger abrupt switches in the optimal action. Existing theo…
Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees
Zuyuan Zhang, Arnob Ghosh, Tian Lan
Making decisions with respect to just the expected returns in Monte Carlo Tree Search (MCTS) cannot account for the potential range of high-risk, adverse outcomes associated with a…
Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks
Zuyuan Zhang, Tian Lan
Monte Carlo Tree Search (MCTS) has proven highly effective in solving complex planning tasks by balancing exploration and exploitation using Upper Confidence Bound for Trees (UCT).…