3 papers
cs.LG2026
Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL
Matthew Zurek, Guy Zamir, Yudong Chen
We study offline reinforcement learning in average-reward MDPs, which presents increased challenges from the perspectives of distribution shift and non-uniform coverage, and has be…
cs.LG2026
Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs
Guy Zamir, Matthew Zurek, Yudong Chen
Online reinforcement learning in infinite-horizon Markov decision processes (MDPs) remains less theoretically and algorithmically developed than its episodic counterpart, with many…
cs.LG2025
Improving Learning to Optimize Using Parameter Symmetries
Guy Zamir, Aryan Dokania, Bo Zhao +1
We analyze a learning-to-optimize (L2O) algorithm that exploits parameter space symmetry to enhance optimization efficiency. Prior work has shown that jointly learning symmetry tra…