5 papers
An Adaptive Method for Contextual Stochastic Multi-armed Bandits with Rewards Generated by a Linear Dynamical System
Jonathan Gornet, Mehdi Hosseinzadeh, Bruno Sinopoli
Online decision-making can be formulated as the popular stochastic multi-armed bandit problem where a learner makes decisions (or takes actions) to maximize cumulative rewards coll…
A Control Theory inspired Exploration Method for a Linear Bandit driven by a Linear Gaussian Dynamical System
Jonathan Gornet, Yilin Mo, Bruno Sinopoli
The paper introduces a linear bandit environment where the reward is the output of a known Linear Gaussian Dynamical System (LGDS). In this environment, we address the fundamental…
HyperController: A Hyperparameter Controller for Fast and Stable Training of Reinforcement Learning Neural Networks
Jonathan Gornet, Yiannis Kantaros, Bruno Sinopoli
We introduce Hyperparameter Controller (HyperController), a computationally efficient algorithm for hyperparameter optimization during training of reinforcement learning neural net…
An Exploration-free Method for a Linear Stochastic Bandit Driven by a Linear Gaussian Dynamical System
Jonathan Gornet, Yilin Mo, Bruno Sinopoli
In stochastic multi-armed bandits, a major problem the learner faces is the trade-off between exploration and exploitation. Recently, exploration-free methods -- methods that commi…
Restless Bandit Problem with Rewards Generated by a Linear Gaussian Dynamical System
Jonathan Gornet, Bruno Sinopoli
Decision-making under uncertainty is a fundamental problem encountered frequently and can be formulated as a stochastic multi-armed bandit problem. In the problem, the learner inte…