4 citations · 4 across the 2 of their papers we have counts for
3 papers
cs.LG2021
Regularized OFU: an Efficient UCB Estimator forNon-linear Contextual Bandit
Yichi Zhou, Shihong Song, Huishuai Zhang +3
Balancing exploration and exploitation (EE) is a fundamental problem in contex-tual bandit. One powerful principle for EE trade-off isOptimism in Face of Uncer-tainty(OFU), in whic…
cs.LG2018
Lazy-CFR: fast and near optimal regret minimization for extensive games with imperfect information
Yichi Zhou, Tongzheng Ren, Jialian Li +2
Counterfactual regret minimization (CFR) is the most popular algorithm on solving two-player zero-sum extensive games with imperfect information and achieves state-of-the-art perfo…
cs.LG2017★ 4 cited
Racing Thompson: an Efficient Algorithm for Thompson Sampling with Non-conjugate Priors
Yichi Zhou, Jun Zhu, Jingwei Zhuo
Thompson sampling has impressive empirical performance for many multi-armed bandit problems. But current algorithms for Thompson sampling only work for the case of conjugate priors…