activity
20172024
most citedAnalysis of Thompson Sampling for Controlling Unknown Linear Diffusion Processes

3 citations · 5 across the 12 of their papers we have counts for

collaborators

17 papers

stat.ML2024

On the Effect of Instability on Learning Continuous-Time Linear Control Systems

Reza Sadeghi Hafshejani, Mohamad Kazem Shirani Fradonbeh

We study the problem of system identification for stochastic continuous-time dynamics, based on a single finite-length state trajectory. We present a method for estimating the poss…

stat.ML2024

Thompson Sampling in Partially Observable Contextual Bandits

Hongju Park, Mohamad Kazem Shirani Faradonbeh

Contextual bandits constitute a classical framework for decision-making under uncertainty. In this setting, the goal is to learn the arms of highest reward subject to contextual in…

cs.LG2022★ 3 cited

Analysis of Thompson Sampling for Controlling Unknown Linear Diffusion Processes

Mohamad Kazem Shirani Faradonbeh, Sadegh Shirani, Mohsen Bayati

Linear diffusion processes serve as canonical continuous-time models for dynamic decision-making under uncertainty. These systems evolve according to drift matrices that specify th…

cs.LG2022

Regret Analysis of Certainty Equivalence Policies in Continuous-Time Linear-Quadratic Systems

Mohamad Kazem Shirani Faradonbeh

This work theoretically studies a ubiquitous reinforcement learning policy for controlling the canonical model of continuous-time stochastic linear-quadratic systems. We show that…

stat.ML2022★ 1 cited

Worst-case Performance of Greedy Policies in Bandits with Imperfect Context Observations

Hongju Park, Mohamad Kazem Shirani Faradonbeh

Contextual bandits are canonical models for sequential decision-making under uncertainty in environments with time-varying components. In this setting, the expected reward of each…

stat.ML2022

Efficient Algorithms for Learning to Control Bandits with Unobserved Contexts

Hongju Park, Mohamad Kazem Shirani Faradonbeh

Contextual bandits are widely-used in the study of learning-based control policies for finite action spaces. While the problem is well-studied for bandits with perfectly observed c…