3 papers
cs.LG2024
Highway Value Iteration Networks
Yuhui Wang, Weida Li, Francesco Faccio +2
Value iteration networks (VINs) enable end-to-end learning for planning tasks by employing a differentiable "planning module" that approximates the value iteration algorithm. Howev…
cs.LG2024
Highway Reinforcement Learning
Yuhui Wang, Miroslav Strupl, Francesco Faccio +5
Learning from multi-step off-policy data collected by a set of policies is a core problem of reinforcement learning (RL). Approaches based on importance sampling (IS) often suffer…
cs.LG2021
Greedy-Step Off-Policy Reinforcement Learning
Yuhui Wang, Qingyuan Wu, Pengcheng He +1
Most of the policy evaluation algorithms are based on the theories of Bellman Expectation and Optimality Equation, which derive two popular approaches - Policy Iteration (PI) and V…