12 papers
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence
Runlong Zhou, Zihan Zhang, Maryam Fazel +1
We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with states, actions, horizon , and per-trajectory total…
Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning
Zijun Chen, Zihan Zhang
We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episod…
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
Ruizhe Shi, Minhak Song, Runlong Zhou +3
We present a fine-grained theoretical analysis of the performance gap between two-stage reinforcement learning from human feedback~(RLHF) and direct preference optimization~(DPO).…
Frozen Policy Iteration: Computationally Efficient RL under Linear Realizability for Deterministic Dynamics
Yijing Ke, Zihan Zhang, Ruosong Wang
We study computationally and statistically efficient reinforcement learning under the linear realizability assumption, where any policy's -function is linear in a given s…
Safe-SDL:Establishing Safety Boundaries and Control Mechanisms for AI-Driven Self-Driving Laboratories
Zihan Zhang, Haohui Que, Junhan Chang +3
The emergence of Self-Driving Laboratories (SDLs) transforms scientific discovery methodology by integrating AI with robotic automation to create closed-loop experimental systems c…
TodoEvolve: Learning to Architect Agent Planning Systems
Jiaxi Liu, Yanzuo Jiang, Guibin Zhang +5
Planning has become a central capability for contemporary agent systems in navigating complex, long-horizon tasks, yet existing approaches predominantly rely on fixed, hand-crafted…