activity
20202022
most citedOnline Learning for Stochastic Shortest Path Model via Posterior Sampling

6 citations · 10 across the 5 of their papers we have counts for

collaborators

8 papers

cs.LG2022

Near-Optimal Goal-Oriented Reinforcement Learning in Non-Stationary Environments

Liyu Chen, Haipeng Luo

We initiate the study of dynamic regret minimization for goal-oriented reinforcement learning modeled by a non-stationary stochastic shortest path problem with changing cost and tr…

cs.LG20221 cited

Policy Learning and Evaluation with Randomized Quasi-Monte Carlo

Sebastien M. R. Arnold, Pierre L'Ecuyer, Liyu Chen +2

Reinforcement learning constantly deals with hard integrals, for example when computing expectations in policy evaluation and policy iteration. These integrals are rarely analytica…

cs.LG2022

Policy Optimization for Stochastic Shortest Path

Liyu Chen, Haipeng Luo, Aviv Rosenberg

Policy optimization is among the most popular and successful reinforcement learning algorithms, and there is increasing interest in understanding its theoretical guarantees. In thi…

cs.LG20223 cited

Learning Infinite-Horizon Average-Reward Markov Decision Processes with Constraints

Liyu Chen, Rahul Jain, Haipeng Luo

We study regret minimization for infinite-horizon average-reward Markov Decision Processes (MDPs) under cost constraints. We start by designing a policy optimization algorithm with…

cs.LG20216 cited

Online Learning for Stochastic Shortest Path Model via Posterior Sampling

Mehdi Jafarnia-Jahromi, Liyu Chen, Rahul Jain +1

We consider the problem of online reinforcement learning for the Stochastic Shortest Path (SSP) problem modeled as an unknown MDP with an absorbing state. We propose PSRL-SSP, a si…

cs.LG2021

Finding the Stochastic Shortest Path with Low Regret: The Adversarial Cost and Unknown Transition Case

Liyu Chen, Haipeng Luo

We make significant progress toward the stochastic shortest path problem with adversarial costs and unknown transition. Specifically, we develop algorithms that achieve $\widetilde…