2 papers
cs.LG2026
Dreaming Smoothly and Sample Efficiently with Gradient Penalized Latent Dynamics
Romil V. Sonigra, P. R. Kumar
Model-based reinforcement learning improves sample efficiency by learning a world model. However, existing latent world models such as DreamerV3 do not explicitly enforce local smo…
cs.LG2023
Value-Biased Maximum Likelihood Estimation for Model-based Reinforcement Learning in Discounted Linear MDPs
Yu-Heng Hung, Ping-Chun Hsieh, Akshay Mete +1
We consider the infinite-horizon linear Markov Decision Processes (MDPs), where the transition probabilities of the dynamic model can be linearly parameterized with the help of a p…