activity
20242026
collaborators

12 papers

cs.LG2026

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

Runlong Zhou, Zihan Zhang, Maryam Fazel +1

We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with states, actions, horizon , and per-trajectory total…

cs.LG2026

Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning

Zijun Chen, Zihan Zhang

We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episod…

cs.LG2026

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

Ruizhe Shi, Minhak Song, Runlong Zhou +3

We present a fine-grained theoretical analysis of the performance gap between two-stage reinforcement learning from human feedback~(RLHF) and direct preference optimization~(DPO).…

cs.LG2026

Frozen Policy Iteration: Computationally Efficient RL under Linear Realizability for Deterministic Dynamics

Yijing Ke, Zihan Zhang, Ruosong Wang

We study computationally and statistically efficient reinforcement learning under the linear realizability assumption, where any policy's -function is linear in a given s…

cs.RO2026

Safe-SDL:Establishing Safety Boundaries and Control Mechanisms for AI-Driven Self-Driving Laboratories

Zihan Zhang, Haohui Que, Junhan Chang +3

The emergence of Self-Driving Laboratories (SDLs) transforms scientific discovery methodology by integrating AI with robotic automation to create closed-loop experimental systems c…

cs.CL2026

TodoEvolve: Learning to Architect Agent Planning Systems

Jiaxi Liu, Yanzuo Jiang, Guibin Zhang +5

Planning has become a central capability for contemporary agent systems in navigating complex, long-horizon tasks, yet existing approaches predominantly rely on fixed, hand-crafted…