6 papers
Finite-Time Convergence of Distributionally Robust Q-Learning with Linear Function Approximation
Saptarshi Mandal, Yashaswini Murthy, R. Srikant
Distributionally robust reinforcement learning (DRRL) seeks policies that perform well when the deployment transition model differs from the nominal model generating the data. Most…
On the Gaussian Limit of the Output of IIR Filters
Yashaswini Murthy, Bassam Bamieh, R. Srikant
We study the asymptotic distribution of the output of a stable Linear Time-Invariant (LTI) system driven by a non-Gaussian stochastic input. Motivated by longstanding heuristics in…
Convergence of Natural Policy Gradient for a Family of Infinite-State Queueing MDPs
Isaac Grosof, Siva Theja Maguluri, R. Srikant
A wide variety of queueing systems can be naturally modeled as infinite-state Markov Decision Processes (MDPs). In the reinforcement learning (RL) context, a variety of algorithms…
Reinforcement Learning with Segment Feedback
Yihan Du, Anna Winnicki, Gal Dalal +2
Standard reinforcement learning (RL) assumes that an agent can observe a reward for each state-action pair. However, in practical applications, it is often difficult and costly to…
Performance of NPG in Countable State-Space Average-Cost RL
Yashaswini Murthy, Isaac Grosof, Siva Theja Maguluri +1
We consider policy optimization methods in reinforcement learning settings where the state space is arbitrarily large, or even countably infinite. The motivation arises from contro…
Exploration-Driven Policy Optimization in RLHF: Theoretical Insights on Efficient Data Utilization
Yihan Du, Anna Winnicki, Gal Dalal +2
Reinforcement Learning from Human Feedback (RLHF) has achieved impressive empirical successes while relying on a small amount of human feedback. However, there is limited theoretic…