2 papers
cs.LG2025
Harnessing Data from Clustered LQR Systems: Personalized and Collaborative Policy Optimization
Vinay Kanakeri, Shivam Bajaj, Ashwin Verma +2
It is known that reinforcement learning (RL) is data-hungry. To improve sample-efficiency of RL, it has been proposed that the learning algorithm utilize data from 'approximately s…
math.OC2025
Power-Constrained Policy Gradient Methods for LQR
Ashwin Verma, Aritra Mitra, Lintao Ye +1
Consider a discrete-time Linear Quadratic Regulator (LQR) problem solved using policy gradient descent when the system matrices are unknown. The gradient is transmitted across a no…