10 citations · 19 across the 9 of their papers we have counts for
9 papers
Non-Stationary Contextual Bandit Learning via Neural Predictive Ensemble Sampling
Zheqing Zhu, Yueyang Liu, Xu Kuang +1
Real-world applications of contextual bandits often exhibit non-stationarity due to seasonality, serendipity, and evolving social trends. While a number of non-stationary contextua…
Learning-based Two-tiered Online Optimization of Region-wide Datacenter Resource Allocation
Chang-Lin Chen, Hanhan Zhou, Jiayu Chen +8
Online optimization of resource management for large-scale data centers and infrastructures to meet dynamic capacity reservation demands and various practical constraints (e.g., fe…
Scalable Neural Contextual Bandit for Recommender Systems
Zheqing Zhu, Benjamin Van Roy
High-quality recommender systems ought to deliver both innovative and relevant content through effective and exploratory interactions with users. Yet, supervised learning-based neu…
IQL-TD-MPC: Implicit Q-Learning for Hierarchical Model Predictive Control
Rohan Chitnis, Yingchen Xu, Bobak Hashemi +4
Model-based reinforcement learning (RL) has shown great promise due to its sample efficiency, but still struggles with long-horizon sparse-reward tasks, especially in offline setti…
Optimizing Long-term Value for Auction-Based Recommender Systems via On-Policy Reinforcement Learning
Ruiyang Xu, Jalaj Bhandari, Dmytro Korenkevych +4
Auction-based recommender systems are prevalent in online advertising platforms, but they are typically optimized to allocate recommendation slots based on immediate expected retur…
Evaluating Online Bandit Exploration In Large-Scale Recommender System
Hongbo Guo, Ruben Naeff, Alex Nikulkov +1
Bandit learning has been an increasingly popular design choice for recommender system. Despite the strong interest in bandit learning from the community, there remains multiple bot…