4 papers
A Modularized Framework for Piecewise-Stationary Restless Bandits
Kuan-Ta Li, Chia-Chun Lin, Ping-Chun Hsieh +1
We study the piecewise-stationary restless multi-armed bandit (PS-RMAB) problem, where each arm evolves as a Markov chain but \emph{mean rewards may change across unknown segments}…
Cross-Domain Policy Optimization via Bellman Consistency and Hybrid Critics
Ming-Hong Chen, Kuan-Chen Pan, You-De Huang +2
Cross-domain reinforcement learning (CDRL) is meant to improve the data efficiency of RL by leveraging the data samples collected from a source domain to facilitate the learning in…
Non-Stationary Restless Multi-Armed Bandits with Provable Guarantee
Yu-Heng Hung, Ping-Chun Hsieh, Kai Wang
Online restless multi-armed bandits (RMABs) typically assume that each arm follows a stationary Markov Decision Process (MDP) with fixed state transitions and rewards. However, in…
BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RL
Yu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin +3
Bayesian optimization (BO) offers an efficient pipeline for optimizing black-box functions with the help of a Gaussian process prior and an acquisition function (AF). Recently, in…