1 paper
Han Zhong, Zhongren Chen, Zhuoran Yang +2
We study episodic reinforcement learning (RL) in non-stationary linear kernel Markov decision processes (MDPs). In this setting, both the reward function and the transition kernel…