2 papers
cs.LG2026
PAC-Bayesian Reinforcement Learning Trains Generalizable Policies
Abdelkrim Zitouni, Mehdi Hennequin, Juba Agoun +3
We derive a novel PAC-Bayesian generalization bound for reinforcement learning that explicitly accounts for Markov dependencies in the data, through the chain's mixing time. This c…
cs.LG2025
Semi-pessimistic Reinforcement Learning
Jin Zhu, Xin Zhou, Jiaang Yao +5
Offline reinforcement learning (RL) aims to learn an optimal policy from pre-collected data. However, it faces challenges of distributional shift, where the learned policy may enco…