2 papers
cs.LG2024
Model Selection for Average Reward RL with Application to Utility Maximization in Repeated Games
Alireza Masoumian, James R. Wright
In standard RL, a learner attempts to learn an optimal policy for a Markov Decision Process whose structure (e.g. state space) is known. In online model selection, a learner attemp…
cs.LG2021
Sequential Estimation under Multiple Resources: a Bandit Point of View
Alireza Masoumian, Shayan Kiyani, Mohammad Hossein Yassaee
The problem of Sequential Estimation under Multiple Resources (SEMR) is defined in a federated setting. SEMR could be considered as the intersection of statistical estimation and b…