32 citations · 72 across the 8 of their papers we have counts for
13 papers
Pretraining in Deep Reinforcement Learning: A Survey
Zhihui Xie, Zichuan Lin, Junyou Li +2
The past few years have seen rapid progress in combining reinforcement learning (RL) with deep learning. Various breakthroughs ranging from games to robotics have spurred the inter…
Hierarchical Conversational Preference Elicitation with Bandit Feedback
Jinhang Zuo, Songwen Hu, Tong Yu +3
The recent advances of conversational recommendations provide a promising way to efficiently elicit users' preferences via conversational interactions. To achieve this, the recomme…
Differentially Private Temporal Difference Learning with Stochastic Nonconvex-Strongly-Concave Optimization
Canzhe Zhao, Yanjie Ze, Jing Dong +2
Temporal difference (TD) learning is a widely used method to evaluate policies in reinforcement learning. While many TD learning methods have been developed in recent years, little…
Conservative Contextual Combinatorial Cascading Bandit
Kun Wang, Canzhe Zhao, Shuai Li +1
Conservative mechanism is a desirable property in decision-making problems which balance the tradeoff between the exploration and exploitation. We propose the novel \emph{conservat…
An Adversarial Imitation Click Model for Information Retrieval
Xinyi Dai, Jianghao Lin, Weinan Zhang +7
Modern information retrieval systems, including web search, ads placement, and recommender systems, typically rely on learning from user feedback. Click models, which study how use…
Online Influence Maximization under Linear Threshold Model
Shuai Li, Fang Kong, Kejie Tang +2
Online influence maximization (OIM) is a popular problem in social networks to learn influence propagation model parameters and maximize the influence spread at the same time. Most…