activity
20182022
most citedOptimize Neural Fictitious Self-Play in Regret Minimization Thinking

4 citations · 7 across the 7 of their papers we have counts for

collaborators

11 papers

cs.CV2022★ 1 cited

NeIF: Representing General Reflectance as Neural Intrinsics Fields for Uncalibrated Photometric Stereo

Zongrui Li, Qian Zheng, Feishi Wang +3

Uncalibrated photometric stereo (UPS) is challenging due to the inherent ambiguity brought by unknown light. Existing solutions alleviate the ambiguity by either explicitly associa…

cs.LG2021

Thompson Sampling for Unimodal Bandits

Long Yang, Zhao Li, Zehong Hu +4

In this paper, we propose a Thompson Sampling algorithm for \emph{unimodal} bandits, where the expected reward is unimodal over the partially ordered arms. To exploit the unimodal…

cs.AI2021★ 4 cited

Optimize Neural Fictitious Self-Play in Regret Minimization Thinking

Yuxuan Chen, Li Zhang, Shijian Li +1

Optimization of deep learning algorithms to approach Nash Equilibrium remains a significant problem in imperfect information games, e.g. StarCraft and poker. Neural Fictitious Self…

cs.LG2020

On Convergence of Gradient Expected Sarsa()

Long Yang, Gang Zheng, Yu Zhang +3

We study the convergence of with linear function approximation. We show that applying the off-line estimate (multi-step bootstrapping) to $\mathtt{Expe…

cs.LG2020

Sample Complexity of Policy Gradient Finding Second-Order Stationary Points

Long Yang, Qian Zheng, Gang Pan

The goal of policy-based reinforcement learning (RL) is to search the maximal point of its objective. However, due to the inherent non-concavity of its objective, convergence to a…

cs.LG2019

Gradient Q: A Unified Algorithm with Function Approximation for Reinforcement Learning

Long Yang, Yu Zhang, Qian Zheng +2

Full-sampling (e.g., Q-learning) and pure-expectation (e.g., Expected Sarsa) algorithms are efficient and frequently used techniques in reinforcement learning. Q is the firs…