collaborators

6 papers

cs.LG2026

Cochain Perspectives on Temporal-Difference Signals for Learning Beyond Markov Dynamics

Zuyuan Zhang, Sizhe Tang, Tian Lan

Non-Markovian dynamics are commonly found in real-world environments due to long-range dependencies, partial observability, and memory effects. The Bellman equation that is the cen…

cs.AI2026

Structuring Value Representations via Geometric Coherence in Markov Decision Processes

Zuyuan Zhang, Zeyu Fang, Tian Lan

Geometric properties can be leveraged to stabilize and speed reinforcement learning. Existing examples include encoding symmetry structure, geometry-aware data augmentation, and en…

cs.LG2026

Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning

Zeyu Fang, Zuyuan Zhang, Mahdi Imani +1

Model-based offline reinforcement learning is brittle under distribution shift: policy improvement drives rollouts into state--action regions weakly supported by the dataset, where…

cs.LG2026

Geometry of Drifting MDPs with Path-Integral Stability Certificates

Zuyuan Zhang, Mahdi Imani, Tian Lan

Real-world reinforcement learning is often \emph{nonstationary}: rewards and dynamics drift, accelerate, oscillate, and trigger abrupt switches in the optimal action. Existing theo…

cs.LG2025

Tail-Risk-Safe Monte Carlo Tree Search under PAC-Level Guarantees

Zuyuan Zhang, Arnob Ghosh, Tian Lan

Making decisions with respect to just the expected returns in Monte Carlo Tree Search (MCTS) cannot account for the potential range of high-risk, adverse outcomes associated with a…

cs.AI2025

Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks

Zuyuan Zhang, Tian Lan

Monte Carlo Tree Search (MCTS) has proven highly effective in solving complex planning tasks by balancing exploration and exploitation using Upper Confidence Bound for Trees (UCT).…