2 papers
cs.LG2026
Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification
Anders Jonsson, Emilie Kaufmann, Gianmarco Tedeschi +1
We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. HBPI-UCRL exploits the fact t…
cs.LG2026
Learning The Minimum Action Distance
Lorenzo Steccanella, Joshua B. Evans, Ãzgür ÅimÅek +1
This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectories, requiring neither reward signals nor t…