6 citations · 19 across the 13 of their papers we have counts for
7 papers · 1 filter
Stochastic Graph Bandit Learning with Side-Observations
Xueping Gong, Jiheng Zhang
In this paper, we investigate the stochastic contextual bandit with general function space and graph feedback. We propose an algorithm that addresses this problem by adapting to bo…
Efficient Transfer Learning via Causal Bounds
Xueping Gong, Wei You, Jiheng Zhang
Transfer learning seeks to accelerate sequential decision-making by leveraging offline data from related agents. However, data from heterogeneous sources that differ in observed fe…
Debiasing Recommendation by Learning Identifiable Latent Confounders
Qing Zhang, Xiaoying Zhang, Yang Liu +4
Recommendation systems aim to predict users' feedback on items not exposed to them. Confounding bias arises due to the presence of unmeasured variables (e.g., the socio-economic st…
Optimal Contextual Bandits with Knapsacks under Realizability via Regression Oracles
Yuxuan Han, Jialin Zeng, Yang Wang +2
We study the stochastic contextual bandit with knapsacks (CBwK) problem, where each action, taken upon a context, not only leads to a random reward but also costs a random resource…
Dual Instrumental Method for Confounded Kernelized Bandits
Xueping Gong, Jiheng Zhang
The contextual bandit problem is a theoretically justified framework with wide applications in various fields. While the previous study on this problem usually requires independenc…
Distributionally Robust Offline Reinforcement Learning with Linear Function Approximation
Xiaoteng Ma, Zhipeng Liang, Jose Blanchet +5
Among the reasons hindering reinforcement learning (RL) applications to real-world problems, two factors are critical: limited data and the mismatch between the testing environment…