activity
20172020
most citedNeural Contextual Bandits with Deep Representation and Shallow Exploration

18 citations · 61 across the 5 of their papers we have counts for

collaborators

14 papers

cs.LG202018 cited

Neural Contextual Bandits with Deep Representation and Shallow Exploration

Pan Xu, Zheng Wen, Handong Zhao +1

We study a general class of contextual bandits, where each context-action pair is associated with a raw feature vector, but the reward generating function is unknown. We propose a…

cs.LG2020

Faster Convergence of Stochastic Gradient Langevin Dynamics for Non-Log-Concave Sampling

Difan Zou, Pan Xu, Quanquan Gu

We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our app…

cs.LG2020

MOTS: Minimax Optimal Thompson Sampling

Tianyuan Jin, Pan Xu, Jieming Shi +2

Thompson sampling is one of the most widely used algorithms for many online decision problems, due to its simplicity in implementation and superior empirical performance over other…

cs.LG2020

Double Explore-then-Commit: Asymptotic Optimality and Beyond

Tianyuan Jin, Pan Xu, Xiaokui Xiao +1

We study the multi-armed bandit problem with subgaussian rewards. The explore-then-commit (ETC) strategy, which consists of an exploration phase followed by an exploitation phase,…

cs.LG2019

Rank Aggregation via Heterogeneous Thurstone Preference Models

Tao Jin, Pan Xu, Quanquan Gu +1

We propose the Heterogeneous Thurstone Model (HTM) for aggregating ranked data, which can take the accuracy levels of different users into account. By allowing different noise dist…

cs.LG201915 cited

A Finite-Time Analysis of Q-Learning with Neural Network Function Approximation

Pan Xu, Quanquan Gu

Q-learning with neural network function approximation (neural Q-learning for short) is among the most prevalent deep reinforcement learning algorithms. Despite its empirical succes…