4 papers · 1 filter
Implicit Q-Learning and SARSA: Liberating Policy Control from Step-Size Calibration
Hwanwoo Kim, Eric Laber
Q-learning and SARSA are foundational reinforcement learning algorithms whose practical success depends critically on step-size calibration. Step-sizes that are too large can cause…
Implicit Updates for Average-Reward Temporal Difference Learning
Hwanwoo Kim, Dongkyu Derek Cho, Eric Laber
Temporal difference (TD) learning is a cornerstone of reinforcement learning. In the average-reward setting, standard TD() is highly sensitive to the choice of step-size and th…
Exploiting Concavity Information in Gaussian Process Contextual Bandit Optimization
Kevin Li, Eric Laber
The contextual bandit framework is widely used to solve sequential optimization problems where the reward of each decision depends on auxiliary context variables. In settings such…
Empirical Bound Information-Directed Sampling for Norm-Agnostic Bandits
Piotr M. Suder, Eric Laber
Information-directed sampling (IDS) is a powerful framework for solving bandit problems which has shown strong results in both Bayesian and frequentist settings. However, frequenti…