40 citations · 125 across the 25 of their papers we have counts for
8 papers · 1 filter
Pretraining in Deep Reinforcement Learning: A Survey
Zhihui Xie, Zichuan Lin, Junyou Li +2
The past few years have seen rapid progress in combining reinforcement learning (RL) with deep learning. Various breakthroughs ranging from games to robotics have spurred the inter…
Hierarchical Conversational Preference Elicitation with Bandit Feedback
Jinhang Zuo, Songwen Hu, Tong Yu +3
The recent advances of conversational recommendations provide a promising way to efficiently elicit users' preferences via conversational interactions. To achieve this, the recomme…
Federated Online Clustering of Bandits
Xutong Liu, Haoru Zhao, Tong Yu +2
Contextual multi-armed bandit (MAB) is an important sequential decision-making problem in recommendation systems. A line of works, called the clustering of bandits (CLUB), utilize…
Comparison-based Conversational Recommender System with Relative Bandit Feedback
Zhihui Xie, Tong Yu, Canzhe Zhao +1
With the recent advances of conversational recommendations, the recommender system is able to actively and dynamically elicit user preference via conversational interactions. To ac…
Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model
Cheng Chen, Canzhe Zhao, Shuai Li
Online learning to rank (OLTR) interactively learns to choose lists of items from a large collection based on certain click models that describe users' click behaviors. Most recent…
A Graph-Enhanced Click Model for Web Search
Jianghao Lin, Weiwen Liu, Xinyi Dai +6
To better exploit search logs and model users' behavior patterns, numerous click models are proposed to extract users' implicit interaction feedback. Most traditional click models…