3 papers
cs.LG2026
Tighter Regret Bounds for Contextual Action-Set Reinforcement Learning
Zijun Chen, Zihan Zhang
We study episodic reinforcement learning with fixed reward and transition functions, but with episode-dependent admissible action sets that are observed at the start of each episod…
cs.LG2026
Sample Complexity of Distributionally Robust Average-Reward Reinforcement Learning
Zijun Chen, Shengbo Wang, Nian Si
Motivated by practical applications where stable long-term performance is critical-such as robotics, operations research, and healthcare-we study the problem of distributionally ro…
cs.LG2026
Achieving Dependence for Average-Reward Q-Learning with a New Contraction Principle
Zijun Chen, Zaiwei Chen, Nian Si +1
We present the convergence rates of synchronous and asynchronous Q-learning for average-reward Markov decision processes, where the absence of contraction poses a fundamental chall…