4 citations · 8 across the 6 of their papers we have counts for
3 papers
cs.LG2023★ 1 cited
Variance-Dependent Regret Bounds for Linear Bandits and Reinforcement Learning: Adaptivity and Computational Efficiency
Heyang Zhao, Jiafan He, Dongruo Zhou +2
Recently, several studies (Zhou et al., 2021a; Zhang et al., 2021b; Kim et al., 2021; Zhou and Gu, 2022) have provided variance-dependent regret bounds for linear contextual bandit…
cs.LG2022
Learning Two-Player Mixture Markov Games: Kernel Function Approximation and Correlated Equilibrium
Chris Junchi Li, Dongruo Zhou, Quanquan Gu +1
We consider learning Nash equilibria in two-player zero-sum Markov Games with nonlinear function approximation, where the action-value function is approximated by a function in a R…
cs.LG2022★ 4 cited
Learning Neural Contextual Bandits Through Perturbed Rewards
Yiling Jia, Weitong Zhang, Dongruo Zhou +2
Thanks to the power of representation learning, neural contextual bandit algorithms demonstrate remarkable performance improvement against their classical counterparts. But because…