7 papers
Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes
Asha Barua, Sajad Khodadadian
Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy…
Tail Distribution of Regret in Optimistic Reinforcement Learning
Sajad Khodadadian, Mehrdad Moharrami
We derive instance-dependent tail bounds for the regret of optimism-based reinforcement learning in finite-horizon tabular Markov decision processes with unknown transition dynamic…
Leveraging High-Fidelity Digital Models and Reinforcement Learning for Mission Engineering: A Case Study of Aerial Firefighting Under Perfect Information
İbrahim OÄuz Ãetinkaya, Sajad Khodadadian, Taylan G. Topcu
As systems engineering (SE) objectives evolve from design and operation of monolithic systems to complex System of Systems (SoS), the discipline of Mission Engineering (ME) has eme…
Optimistic Reinforcement Learning with Quantile Objectives
Mohammad Alipour-Vaezi, Huaiyang Zhong, Kwok-Leung Tsui +1
Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective funct…
A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging
Sajad Khodadadian, Martin Zubeldia
Polyak-Ruppert averaging is a widely used technique to achieve the optimal asymptotic variance of stochastic approximation (SA) algorithms, yet its high-probability performance gua…
Tight Finite Time Bounds of Two-Time-Scale Linear Stochastic Approximation with Markovian Noise
Shaan Ul Haque, Sajad Khodadadian, Siva Theja Maguluri
Stochastic approximation (SA) is an iterative algorithm for finding the fixed point of an operator using noisy samples and widely used in optimization and Reinforcement Learning (R…