7 papers
Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation
Saptarshi Mandal, Yashaswini Murthy, R. Srikant
Distributionally robust reinforcement learning seeks policies that remain effective when the deployment environment differs from the one that generated the training data. We study…
On the Gaussian Limit of the Output of IIR Filters
Yashaswini Murthy, Bassam Bamieh, R. Srikant
We study the asymptotic distribution of the output of a stable Linear Time-Invariant (LTI) system driven by a non-Gaussian stochastic input. Motivated by longstanding heuristics in…
Reinforcement Learning with Segment Feedback
Yihan Du, Anna Winnicki, Gal Dalal +2
Standard reinforcement learning (RL) assumes that an agent can observe a reward for each state-action pair. However, in practical applications, it is often difficult and costly to…
Performance of NPG in Countable State-Space Average-Cost RL
Yashaswini Murthy, Isaac Grosof, Siva Theja Maguluri +1
We consider policy optimization methods in reinforcement learning settings where the state space is arbitrarily large, or even countably infinite. The motivation arises from contro…
On the Global Convergence of Policy Gradient in Average Reward Markov Decision Processes
Navdeep Kumar, Yashaswini Murthy, Itai Shufaro +3
We present the first finite time global convergence analysis of policy gradient in the context of infinite horizon average reward Markov decision processes (MDPs). Specifically, we…
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
Yihan Du, Anna Winnicki, Gal Dalal +2
Reinforcement Learning from Human Feedback (RLHF) has achieved impressive empirical successes while relying on a small amount of human feedback. However, there is limited theoretic…