1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2024
Dual Approximation Policy Optimization
Zhihan Xiong, Maryam Fazel, Lin Xiao
We propose Dual Approximation Policy Optimization (DAPO), a framework that incorporates general function approximation into policy mirror descent methods. In contrast to the popula…
cs.LG2021★ 1 cited
Selective Sampling for Online Best-arm Identification
Romain Camilleri, Zhihan Xiong, Maryam Fazel +2
This work considers the problem of selective-sampling for best-arm identification. Given a set of potential options , a learner aims to compute with…
cs.LG2019
Parameterized Indexed Value Function for Efficient Exploration in Reinforcement Learning
Tian Tan, Zhihan Xiong, Vikranth R. Dwaracherla
It is well known that quantifying uncertainty in the action-value estimates is crucial for efficient exploration in reinforcement learning. Ensemble sampling offers a relatively co…