activity
20212026
most citedInformation Theoretic Measures for Fairness-aware Feature Selection

5 citations · 5 across the 6 of their papers we have counts for

collaborators

11 papers

cs.LG2026

Risk-Sensitive Reinforcement Learning with Smoothed Quantile Objectives

Mohammad Alipour-Vaezi, Huaiyang Zhong, Sajad Khodadadian

Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective funct…

cs.LG2026

Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes

Asha Barua, Sajad Khodadadian

Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy…

cs.CY2025

Leveraging High-Fidelity Digital Models and Reinforcement Learning for Mission Engineering: A Case Study of Aerial Firefighting Under Perfect Information

İbrahim Oğuz Çetinkaya, Sajad Khodadadian, Taylan G. Topcu

As systems engineering (SE) objectives evolve from design and operation of monolithic systems to complex System of Systems (SoS), the discipline of Mission Engineering (ME) has eme…

cs.LG2025

Optimistic Reinforcement Learning with Quantile Objectives

Mohammad Alipour-Vaezi, Huaiyang Zhong, Kwok-Leung Tsui +1

Reinforcement Learning (RL) has achieved tremendous success in recent years. However, the classical foundations of RL do not account for the risk sensitivity of the objective funct…

cs.LG2025

Tail Distribution of Regret in Optimistic Reinforcement Learning

Sajad Khodadadian, Mehrdad Moharrami

We derive instance-dependent tail bounds for the regret of optimism-based reinforcement learning in finite-horizon tabular Markov decision processes with unknown transition dynamic…

stat.ML2025

A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak Averaging

Sajad Khodadadian, Martin Zubeldia

Polyak-Ruppert averaging is a widely used technique to achieve the optimal asymptotic variance of stochastic approximation (SA) algorithms, yet its high-probability performance gua…