activity
20122021
most citedEmphatic Temporal-Difference Learning

21 citations · 49 across the 5 of their papers we have counts for

collaborators

9 papers

math.OC20212 cited

Average-Cost Optimality Results for Borel-Space Markov Decision Processes with Universally Measurable Policies

Huizhen Yu

We consider discrete-time Markov Decision Processes with Borel state and action spaces and universally measurable policies. For several long-run average cost criteria, we establish…

math.OC2020

Average Cost Optimality Inequality for Markov Decision Processes with Borel Spaces and Universally Measurable Policies

Huizhen Yu

We consider average-cost Markov decision processes (MDPs) with Borel state and action spaces and universally measurable policies. For the nonnegative cost model and an unbounded co…

math.OC2019

On the Minimum Pair Approach for Average-Cost Markov Decision Processes with Countable Discrete Action Spaces and Strictly Unbounded Costs

Huizhen Yu

We consider average-cost Markov decision processes (MDPs) with Borel state spaces, countable, discrete action spaces, and strictly unbounded one-stage costs. For the minimum pair a…

math.OC20194 cited

On Markov Decision Processes with Borel Spaces and an Average Cost Criterion

Huizhen Yu

We consider average-cost Markov decision processes (MDPs) with Borel state and action spaces and universally measurable policies. For the nonnegative cost model and an unbounded co…

cs.LG2018

Two geometric input transformation methods for fast online reinforcement learning with neural nets

Sina Ghiassian, Huizhen Yu, Banafsheh Rafiee +1

We apply neural nets with ReLU gates in online reinforcement learning. Our goal is to train these networks in an incremental manner, without the computationally expensive experienc…

cs.LG201721 cited

Multi-step Off-policy Learning Without Importance Sampling Ratios

Ashique Rupam Mahmood, Huizhen Yu, Richard S. Sutton

To estimate the value functions of policies from exploratory data, most model-free off-policy algorithms rely on importance sampling, where the use of importance sampling ratios of…